Ship an SDK so AI agents build integrations that work in production

5 min readThe Fern Team
Three stacked windows: an openapi.yml spec, a Fern SDK generated from it, and a Claude Code terminal where the prompt "Charge a customer $10 with Square" produces an await square.payments.create call and "Done! Payment completed on the first run."

TL;DR

Ship an SDK so AI agents building on your API get it right the first time. In our benchmark on a real payments API, 10 of 12 SDK integrations were correct and survived production rate limits, versus 2 of 12 on raw HTTP, and agents spent 77% to 100% fewer tokens fixing their own bugs. Generate yours from your OpenAPI spec with Fern.

AI agents are building bad integrations against your API, and your customers are paying for the tokens it takes to fix them.

Your customers hand integrations to coding agents now. Whatever the agent writes is the code that calls your API in production. Its bugs become your support tickets.

So we measured it. Same agent, same task, same API, with and without an SDK. The API: Square. The SDK: Square's TypeScript SDK, generated by Fern.

Without an SDK, agent-built integrations break under load

Every Sonnet integration built on raw HTTP broke under production rate limits. All six passed the agent's own testing. All six crashed the moment the API started returning 429.

Every Sonnet integration built on Square's SDK survived. Six of six, with no retry code written by the agent.

The raw HTTP failure is worth slowing down on. Five of the six Sonnet sessions hit a 429 during development. Every time, the agent waited a few seconds, re-ran the script, saw it pass, and called the job done. None of the six wrote retry logic. That bug ships, passes code review, and waits for your customer's first busy day.

The test

Square publishes its OpenAPI spec: 253 endpoints, 1,476 schemas. We built a strict mock server from it. It validates request bodies against the spec's schemas, requires idempotency keys, enforces auth, and returns Square's real error format.

The task: find the customer with email ops@globex.com, create an order at the "Downtown" location with two line items, charge the customer's Visa on file, print the result. The traps are the ones real data has. 230 paginated customers. Lookalike emails like ops@globex.co. A second location. An expired card.

Two setups, six runs each, on Claude Sonnet 4.6 and Claude Haiku 4.5. Each agent had a shell, a file-writing tool, and the live mock to test against:

  • Raw HTTP: Square's OpenAPI spec and Node's built-in fetch.
  • Square SDK: the published square package from npm, generated by Fern, with its stock README and reference.

When the agent finished, we ran its script against a fresh server and checked the result: one order, right customer, right location, right line items, paid once with the right card. Then we ran it again against a server that rate-limits every second request.

The results

ModelSetup
Sonnet 4.6Raw HTTP0 of 61.074k342k15295 s
Sonnet 4.6Square SDK6 of 600337k11282 s
Haiku 4.5Raw HTTP2 of 66.2769k1.15M232200 s
Haiku 4.5Square SDK4 of 60.7178k1.08M122107 s

Every number except "Ready" is the average for one session, across the six runs for that setup.

What the SDK changed, compared with raw HTTP:

Model
Sonnet 4.6-100%-100%-2%-26%-14%
Haiku 4.5-89%-77%-6%-47%-47%

With the SDK, Sonnet got it right on the first run in all six sessions. Haiku crashed 89% less often and finished in about half the time.

Fixing bugs is where the tokens go

Every failed run costs your customer twice. Once to read the error, again to re-send the whole conversation while the agent patches it.

Haiku on raw HTTP spent 67% of its tokens repairing its own code: 769k tokens per session. With Square's SDK, repairs dropped to 178k, 77% fewer.

Sonnet on raw HTTP spent 74k tokens on repairs. With the SDK, zero.

The SDK sessions put those tokens into building instead. Once Sonnet had read the docs, building the integration took 20% fewer tokens with the SDK than building and fixing it with raw HTTP.

The raw HTTP mistakes were familiar:

  • No rate-limit handling. 10 of the 12 raw HTTP scripts had no 429 handling at all.
  • Invented endpoints. Haiku called GET /v2/customers/{id}/cards, which doesn't exist.
  • Out-of-range parameters. A page limit above Square's maximum of 100.
  • Wrong payment flow. Paying an order that the payment had already completed.

The SDK closes off most of these. Base URL and auth are set once on the client. Retries are built in. Pagination is an iterator. Request fields are typed, so a wrong one fails at compile time instead of in production.

Bad integrations become the API owner's problem

A raw HTTP integration without retries looks fine until traffic spikes. Then your customer's payments fail, their users complain, and the ticket lands in your queue as "your API is down." Your team spends a day proving it isn't.

Hand-written retries are their own risk. Each script picks its own attempts and delays, and one aggressive loop turns a slowdown on your side into a retry storm from your own customers.

Square's SDK retries 408, 429, and 5xx responses with exponential backoff and jitter, honors Retry-After, and stops after a bounded number of attempts. Square sets that policy once. Every customer on the SDK follows it.

One code path instead of thousands

The raw HTTP scripts ran 131 to 289 lines, and each one took its own approach to auth, pagination, and errors. Multiply that by your customer base and you are supporting thousands of one-off clients.

That variation costs the API owner:

  • Support. Every ticket starts with reverse-engineering the customer's request code.
  • API changes. Hand-written types break the moment you add a field, and you don't know who breaks.
  • Fixes. You can't ship a better retry policy or a new auth flow to code you didn't write.
  • Visibility. Raw requests don't say what sent them.

With the SDK, every customer runs the same tested request path, which you own. A fix ships with a version bump. And every request carries X-Fern-SDK-Name and X-Fern-SDK-Version, so you see exactly which version each customer runs.

Give agents an SDK to find

Agents build from whatever they find first. Make it a typed SDK.

Fern generates SDKs in 9 languages from your OpenAPI spec, with retries, pagination, and typed errors built in, and regenerates them every time your API changes. Square, ElevenLabs, and Cohere ship theirs with Fern.

Book a demo and we'll generate one from your spec.

The Fern Team