# Ship an SDK so AI agents build integrations that work in production **TL;DR:** Ship an SDK so AI agents building on your API get it right the first time. In our benchmark on a real payments API, 10 of 12 SDK integrations were correct and survived production rate limits, versus 2 of 12 on raw HTTP, and agents spent 77% to 100% fewer tokens fixing their own bugs. Generate yours from your OpenAPI spec with Fern. AI agents are building bad integrations against your API, and your customers are paying for the tokens it takes to fix them. Your customers hand integrations to coding agents now. Whatever the agent writes is the code that calls your API in production. Its bugs become your support tickets. So we measured it. Same agent, same task, same API, with and without an SDK. The API: Square. The SDK: Square's TypeScript SDK, generated by Fern. ## Without an SDK, agent-built integrations break under load **Every Sonnet integration built on raw HTTP broke under production rate limits.** All six passed the agent's own testing. All six crashed the moment the API started returning `429`. **Every Sonnet integration built on Square's SDK survived.** Six of six, with no retry code written by the agent. The raw HTTP failure is worth slowing down on. Five of the six Sonnet sessions hit a `429` during development. Every time, the agent waited a few seconds, re-ran the script, saw it pass, and called the job done. None of the six wrote retry logic. That bug ships, passes code review, and waits for your customer's first busy day. ## The test Square publishes its OpenAPI spec: 253 endpoints, 1,476 schemas. We built a strict mock server from it. It validates request bodies against the spec's schemas, requires idempotency keys, enforces auth, and returns Square's real error format. The task: find the customer with email `ops@globex.com`, create an order at the "Downtown" location with two line items, charge the customer's Visa on file, print the result. The traps are the ones real data has. 230 paginated customers. Lookalike emails like `ops@globex.co`. A second location. An expired card. Two setups, six runs each, on Claude Sonnet 4.6 and Claude Haiku 4.5. Each agent had a shell, a file-writing tool, and the live mock to test against: - **Raw HTTP:** Square's OpenAPI spec and Node's built-in `fetch`. - **Square SDK:** the published `square` package from npm, generated by Fern, with its stock README and reference. When the agent finished, we ran its script against a fresh server and checked the result: one order, right customer, right location, right line items, paid once with the right card. Then we ran it again against a server that rate-limits every second request. ## The results | Model | Setup | Ready | Crashes | Fix tokens | Tokens | Lines | Time | | ---------- | ---------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | ----------------------------------------------------------------- | | Sonnet 4.6 | Raw HTTP | 0 of 6 | 1.0 | 74k | 342k | 152 | 95 s | | Sonnet 4.6 | Square SDK | 6 of 6 | 0 | 0 | 337k | 112 | 82 s | | Haiku 4.5 | Raw HTTP | 2 of 6 | 6.2 | 769k | 1.15M | 232 | 200 s | | Haiku 4.5 | Square SDK | 4 of 6 | 0.7 | 178k | 1.08M | 122 | 107 s | Every number except "Ready" is the average for one session, across the six runs for that setup. What the SDK changed, compared with raw HTTP: | Model | Crashes | Fix tokens | Tokens | Lines | Time | | ---------- | ------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | ----------------------------------------------------------------- | | Sonnet 4.6 | -100% | -100% | -2% | -26% | -14% | | Haiku 4.5 | -89% | -77% | -6% | -47% | -47% | With the SDK, Sonnet got it right on the first run in all six sessions. Haiku crashed 89% less often and finished in about half the time. ## Fixing bugs is where the tokens go Every failed run costs your customer twice. Once to read the error, again to re-send the whole conversation while the agent patches it. **Haiku on raw HTTP spent 67% of its tokens repairing its own code: 769k tokens per session.** With Square's SDK, repairs dropped to 178k, 77% fewer. **Sonnet on raw HTTP spent 74k tokens on repairs. With the SDK, zero.** The SDK sessions put those tokens into building instead. Once Sonnet had read the docs, building the integration took 20% fewer tokens with the SDK than building and fixing it with raw HTTP. The raw HTTP mistakes were familiar: - **No rate-limit handling.** 10 of the 12 raw HTTP scripts had no `429` handling at all. - **Invented endpoints.** Haiku called `GET /v2/customers/{id}/cards`, which doesn't exist. - **Out-of-range parameters.** A page `limit` above Square's maximum of 100. - **Wrong payment flow.** Paying an order that the payment had already completed. The SDK closes off most of these. Base URL and auth are set once on the client. Retries are built in. Pagination is an iterator. Request fields are typed, so a wrong one fails at compile time instead of in production. ## Bad integrations become the API owner's problem A raw HTTP integration without retries looks fine until traffic spikes. Then your customer's payments fail, their users complain, and the ticket lands in your queue as "your API is down." Your team spends a day proving it isn't. Hand-written retries are their own risk. Each script picks its own attempts and delays, and one aggressive loop turns a slowdown on your side into a retry storm from your own customers. Square's SDK retries `408`, `429`, and `5xx` responses with exponential backoff and jitter, honors `Retry-After`, and stops after a bounded number of attempts. Square sets that policy once. Every customer on the SDK follows it. ## One code path instead of thousands The raw HTTP scripts ran 131 to 289 lines, and each one took its own approach to auth, pagination, and errors. Multiply that by your customer base and you are supporting thousands of one-off clients. That variation costs the API owner: - **Support.** Every ticket starts with reverse-engineering the customer's request code. - **API changes.** Hand-written types break the moment you add a field, and you don't know who breaks. - **Fixes.** You can't ship a better retry policy or a new auth flow to code you didn't write. - **Visibility.** Raw requests don't say what sent them. With the SDK, every customer runs the same tested request path, which you own. A fix ships with a version bump. And every request carries `X-Fern-SDK-Name` and `X-Fern-SDK-Version`, so you see exactly which version each customer runs. ## Give agents an SDK to find Agents build from whatever they find first. Make it a typed SDK. Fern generates SDKs in 9 languages from your OpenAPI spec, with retries, pagination, and typed errors built in, and regenerates them every time your API changes. Square, ElevenLabs, and Cohere ship theirs with Fern. [Book a demo](/contact) and we'll generate one from your spec.