GoatLabelsGoatLabels

Blog

Testing a Shipping API in CI With Test Keys

How to test a shipping API integration in CI: test keys, fixtures, OpenAPI contract checks, idempotent retries, and webhook replay.

You test a shipping API in CI by splitting the work into three layers: fast tests that run against recorded fixtures, a small set of tests that call the provider with a test key, and webhook handler tests fed by saved payloads. The test key keeps money and real carrier labels out of the pipeline, and the fixtures keep most of the suite fast and independent of anyone's uptime. The rest of this post covers what goes in each layer and how to keep the layers honest as the API changes.

Why does a shipping integration need its own test plan?

A shipping integration needs its own plan because a successful call spends money and produces a physical artifact. A bug in a typical integration writes a bad row. A bug in a label integration buys a label nobody uses, or buys the same label twice, or prints a label with the wrong weight that gets adjusted later.

Three properties make it different from most API work:

  • The write is not free. Each purchase debits a balance, so "just run it again" is not a neutral instruction.
  • The result arrives in two parts. The purchase response comes back on the request, and tracking events arrive later through webhooks, possibly out of order and possibly more than once.
  • The inputs are messy. Addresses, weights, and dimensions come from an ERP or order system that was not designed around carrier validation rules.

A test plan that only checks "the request returned 200" misses all three. The plan below is built around them.

What should run against test keys, and what against fixtures?

Most tests should run against fixtures, and a small number should call the API with a test key. Fixtures make the suite fast and deterministic. Test-key calls prove that your real HTTP client, authentication, and request shape are accepted by the real service.

Layer What it calls Runs when What it proves
Unit Nothing over the network; saved request and response JSON Every commit Your mapping from order to request body, and your parsing of responses and errors
Contract The OpenAPI document, not the API Every commit Your fixtures still match the published schema
Integration The API with a test key Merge to main, or nightly Auth, headers, and a full request and response round trip
Webhook Your own handler with saved payloads Every commit Signature checking, duplicate handling, out-of-order handling

Two rules keep this table true in practice. First, the test key lives in the CI secret store and nowhere else, and the live key is not available to the pipeline at all. A job that cannot read the live key cannot spend real money, whatever the code does. Second, the base URL and key are read from the environment, so the same code path runs in tests and in production with different values.

What a test key returns is provider-specific. Do not assume that a test-mode label, rate, or tracking number behaves like a live one. Read the provider's documentation and write the assumption down next to the test that depends on it.

How do you build fixtures that stay accurate?

You keep fixtures accurate by recording them from real test-key responses and validating them against the provider's schema on every run. Hand-written fixtures drift toward what the developer expected the API to return, which is the opposite of their purpose.

A workable routine:

  1. Record a response from a test-key call for each case you care about: a clean purchase, a validation error, an authentication failure, and a rate-limited response.
  2. Strip secrets and anything that looks like a credential before committing the file.
  3. Store the request body beside the response, so the pair documents one exchange.
  4. In CI, validate each stored request and response against the provider's OpenAPI document. A failure means the schema moved or the fixture was edited by hand.
  5. Re-record on a schedule, for example monthly, and review the diff like any other code change.

Step 4 is the one teams skip, and it is the cheapest. If the provider publishes an OpenAPI spec, you can also generate your client types from it, which turns many schema changes into compile errors instead of production incidents.

Cover the unhappy inputs deliberately. Build fixtures for a missing postal code, a zero weight, an overlong address line, and a country your account cannot ship from. Those are the cases an ERP export will eventually produce.

How do you test retries without buying a label twice?

You test retries by sending the same idempotency key on the retry and asserting that your system ends up with exactly one label. The dangerous case is a timeout: the request reached the server and the purchase happened, but your client never saw the response. Without an idempotency key, the retry is a second purchase.

Here is a worked example. The numbers are hypothetical and only illustrate the arithmetic.

Suppose a nightly job ships 400 orders, and 1 percent of purchase calls time out on the client side. That is 4 timeouts. If the job retries each one without an idempotency key, and all 4 original requests had in fact succeeded, the job buys 404 labels for 400 orders. At an example price of $9.00 per label, that is $36.00 of duplicate spend per night, plus the work of finding and voiding the duplicates. With a key derived from the order, the same 4 retries return the original results and the count stays at 400.

The tests that protect this:

  • Key derivation. Assert that the key is built from stable business data, such as the order number plus the carton number, and not from a timestamp or a random value generated per attempt. A key that changes on retry protects nothing.
  • Simulated timeout. In a unit test, make the HTTP layer throw after "sending" and assert that the retry carries the same key.
  • Real replay. In the integration layer, send the same request twice with the same key using the test key, and assert that your code records one shipment. Check the provider's documentation for exactly what the second response looks like.
  • Rate limits. Feed your client a rate-limited response fixture and assert that it backs off using the rate-limit headers the API returns, and that the eventual retry still carries the original key.

How do you test webhooks in a pipeline with no public URL?

You test webhooks in CI by calling your handler directly with saved payloads, because a CI runner has no address a provider can reach. The handler is a function that takes a body and headers. Nothing about testing it requires a live delivery.

Save a real payload and its signature headers from a test-mode delivery, then write these cases:

  1. A valid payload with a valid signature is accepted and updates the order.
  2. The same payload with one byte changed is rejected. This proves you verify the signature against the raw body, not a re-serialized copy.
  3. The same valid payload delivered twice produces one update. Providers retry, and replay features exist precisely so that events can be sent again.
  4. A "delivered" event that arrives before an "in transit" event does not move the order backward.
  5. A payload with an event type your code does not know is acknowledged and ignored, not turned into an error that triggers endless retries.

Use a signing secret that exists only for tests. If the provider supports replaying past events, add a manual step to your release checklist: deploy to staging, replay a handful of recent test events, and confirm the orders update. That covers the networking and routing that the direct handler tests cannot see.

What belongs in the pipeline, in order?

The pipeline should run the cheap, deterministic checks first and the network-dependent checks last. A reasonable order:

  1. Lint and type-check, including generated client types.
  2. Unit tests against fixtures.
  3. Contract validation of fixtures against the OpenAPI document.
  4. Webhook handler tests.
  5. Integration tests with the test key, on merge or nightly, marked so a provider outage does not block unrelated work.
  6. A guard that fails the build if a live key pattern appears in the repository or the job environment.

Keep the integration set small. Five to ten calls that cover authentication, one purchase, one validation error, and one idempotent replay tell you more than a hundred near-duplicates, and they finish quickly.

Where GoatLabels fits

GoatLabels provides a REST API with live and test keys, so a pipeline can hold a test key and never see the live one. Requests authenticate with an Authorization: Bearer header, and shipment requests accept an Idempotency-Key header for safe retries (see /docs for the endpoints). Webhooks are signed and can be replayed, and responses carry rate-limit headers. An OpenAPI 3.1 document is published (see /docs), which is what you would point contract validation and client generation at. The shipping API overview has the developer details, and you should check /docs for what test keys return before writing assertions against test-mode responses.

The limits matter for planning. There are no SDKs, so you generate a client from the spec or write a thin HTTP wrapper. There are no native ERP, WMS, or OMS connectors beyond the 13 store connectors, so an ERP shipping integration or a commercetools shipping integration is code your team writes and tests, using the approach above. Tracking does not write back into an ERP automatically; your webhook handler does that. There is no SLA, which is one more reason to keep integration tests from blocking unrelated merges.