How We Rebuilt Our API Backend Without Breaking Customer Integrations

September 28, 2026

11 min read

Author: Christos Melas

How we rebuilt the Ayrshare API backend

TL;DR

We built Ayrshare 3, a new backend for the Ayrshare API, and moved every customer onto it one account at a time between July and August 2026. No integration had to change.

  • Routes, responses, error codes and the hostname are unchanged. There is no new API version to adopt.
  • A feature flag inside the old backend moved accounts one at a time, and any account could be moved back within about a minute.
  • July performance work cut p95 latency by up to 88.7% on the busiest endpoints, and Ayrshare 3 was built to hold that speed.
  • Every pull request was reviewed and approved by a person before it merged, including the ones drafted with Claude Code.
  • The new foundation helps us deliver platforms, features and fixes sooner.

At 12:30 UTC on 19 August 2026, I repointed the load balancer for api.ayrshare.com at Ayrshare 3, our new backend. By then, customer traffic had been moving to it one account at a time for a month, routed through the backend that had served the API reliably since 2020. This was the last step, taking that old front door out of the path. Then I watched the error graph.

At 12:32, it showed two 502 gateway errors. At 12:33, it showed forty. The new service had cold-started under the full weight of production traffic, and its instances had not yet caught up. I had the rollback command ready in the other terminal. At 12:34, the graph showed zero, and it stayed there.

Forty-two requests. That was the customer-visible cost of the final switch.

If you build on the Ayrshare API, your integration did not need to change. Routes, responses, error codes and the hostname stayed the same. The only differences are a few improvements to webhook delivery and rate limiting.

That was the whole point of the project, and this is the story of how we got there.

  • Early July: backend performance work ships and cuts latency on the most used endpoints.
  • Week of 20 July: the first customer accounts are routed to Ayrshare 3.
  • 19 August: the load balancer moves to Ayrshare 3.
  • 27 August: the last background job moves, and the migration is complete.

Why We Rebuilt a Working API Backend

The previous backend was doing its job. It had served the API since 2020, it was handling tens of millions of calls a day, and customers were building products on it that we were proud to host. Nothing was on fire.

But a successful API keeps growing. A customer needs TikTok, so you add TikTok. A platform changes its rules, so you add handling for those rules. Years of those decisions produce a codebase that reflects what made the business work. They also produce some very large files.

None of that stopped the API from serving customers. But it made each new change more expensive, and by early 2026, we had another reason to address it: we had started working with AI coding agents.

Some files had become too large to fit in an agent’s context window, limiting how effectively it could work with them. The way we wanted to build had outgrown the shape of the code.

We could have kept refactoring incrementally for a long time. Instead, we chose to invest in a foundation for the next five years while the current one was still healthy enough to carry the load during the switch. You rebuild from a position of strength, or you rebuild in a hurry, and only one of those goes well.

Preserving API Backward Compatibility

The product spec started with a strict rule: the rebuild could not change routes, request bodies, response shapes, error codes, headers or the database.

We kept new features and performance work to a minimum. Anything that could wait, waited. The line that governs the rest: customers must not be able to tell the difference.

That constraint has a consequence people underestimate. The new code had to reproduce the existing system's behavior exactly, including its quirks. When the API returns an error in a particular shape, the new backend returns the same shape, even though we would have designed it differently. We kept a list of rough edges we wanted to smooth and deliberately left them for later. Two systems running side by side against one live database could not afford differences in how they read and wrote shared data.

It takes discipline to port something you know you could improve. But preserving backward compatibility was what later let us move customers one at a time without asking anything of them.

Building the New API Backend for Easier Maintenance

The rebuild took nineteen weeks. In the new codebase, request handling, business logic, platform integrations and database access live in separate layers, with clear rules about what each layer may import.

Only one directory talks to the social platforms. Only one directory talks to the database, and everything it reads or writes is validated against a schema. Files and functions have size limits, enforced on every pull request rather than left to good intentions, so a unit of work stays readable by a person and loadable by an agent. Type checking runs on every pull request.

We built one shared layer of fake platform data that serves both the tests and a sandbox mode that runs the real server end-to-end, because a fake that lives in many places drifts, and a drifted fake does worse than miss bugs. It manufactures evidence that there are none.Much of the rebuild was developed with AI agents. I’m sharing this openly, as I’m writing this article under my name. Most pull requests were drafted with Claude Code and labeled accordingly. Each went through automated review, then a person read and approved it before it merged. A branch rule enforced that process., and every file in the new codebase arrived through that gate.

The judgment calls were made by people, file by file. What to keep, what to defer, when the rule had to bend, what was close enough and what was not. The agents were the tool. The code, the typing, the tests and much of the comparison work were done with them, at a volume we could not have reached otherwise, and not one line of it entered the codebase without a person deciding it should. That has not changed since, and the codebase is shaped so that it keeps working.

Before any customer touched the new system, we audited it. Every REST handler was compared behavior-for-behavior with its counterpart. Most matched exactly. A few had drifted and were corrected. We did not assume parity. We went looking for the ways it was false.

Using Feature Flags to Migrate API Traffic Gradually

The original technical spec described a textbook cutover: record production traffic, replay it against the new backend, compare, then swap the load balancer. We designed the replay harness. We never built it.

The closer we got to a working backend, the less suitable that approach looked. It would move every customer at once, with no way to trial an individual account. Rolling it back would require an infrastructure change under pressure. And recorded requests, however faithfully replayed, would still differ from live customer traffic.

We replaced that plan with a feature flag inside the existing backend.

Every request continued to land on the old backend, which authenticated it as usual. Once it knew which account was calling, it checked the flag. During the rollout, a per-account allowlist determined which accounts moved to Ayrshare 3. A global switch would later route everyone to it.

For an account on the allowlist, the old backend forwarded the request to the new one and returned its response unchanged. The client continued using the same API.

What made it the right mechanism was the shape of its controls rather than any cleverness in the routing. Enrolling an account is adding an identifier to that allowlist. Reverting one is removing it. Neither is a deployment, a release, or an infrastructure change.

The gate is read from configuration, so a change reaches every running instance within about a minute. At any point in the rollout, for one customer or for all of them at once, we could immediately put traffic back on the system that had served them since 2020, with nothing to ship and nothing to wait for.

That is what made a gradual rollout safe rather than merely slow. It validated real production traffic from unmodified integrations, with a per-customer blast radius and a per-customer undo, and it shaped everything that came after.

features

Migrating Background Jobs Without Losing Work

The routing flag moved API requests one account at a time. Background work needed a different approach.

Scheduled posts, webhook delivery, token refresh and analytics collection each run globally. Each group, therefore, had to move from the old backend to the new one as a unit, with a clear owner and any overlap deliberately planned.

If two systems refresh the same token, they can race, leaving one with an invalid token. If neither system owns the job, the work stops.

We wrote a runbook for each group of background work, and the index to those runbooks opens with three rules that governed five weeks of handoffs.

  1. Never both, never neither. At any instant, one system owns a piece of work, and the ambiguous window is measured rather than assumed.
  2. Overlap for idempotent work, gap for destructive work. Where doing something twice is harmless, let the systems overlap rather than risk a gap. Where doing something twice causes harm, accept a short gap and rely on a recovery sweep.
  3. A write is not done until a fresh read confirms it. A flag flip is done when a new read shows the new value, not when the command returns.

Underneath all three sat a fourth idea that decided every close call. A duplicate is visible, so it can be fixed. A silent loss is not. When we had to choose which failure to risk, we chose the one that leaves evidence.

Keeping an API Migration Log for Verification

Every step of the migration was put into a single file in the repository. Date, time, the command, what the verification read returned, and what we did about anything unexpected. Entries are never edited; they are only appended. Editing an entry would falsify the record rather than correct it, and the record is the whole point of the file.

It is the least glamorous artifact from the project and the one I would hand to anyone attempting something similar. When something looked wrong at two in the afternoon, the answer to "what exactly did we do at eleven" was a file, not a search through chat history and hope.

API Performance: Where the Latency Improvements Came From

The API is faster this year, and I want to be clear about where that speed came from. The performance work and the rebuild ran side by side, making them easy to confuse.

In early July, a couple of weeks before the first customer was routed to Ayrshare 3, we shipped a round of backend performance work that cut latency across the most heavily used endpoints, with the largest gains under heavy load. No API changes were required. Every request got faster automatically.

Bar chart of latency reductions from the July 2026 performance work. p95 latency fell 88.7% on GET /profiles, 75.0% on GET /analytics, 63.1% on GET /post, 59.4% on GET /messages and 48.6% on GET /user.
Faster before the rebuild landed: the baseline Ayrshare 3 had to hold.
EndpointFaster averageFaster p95
GET /profiles71.6%88.7%
GET /analytics59.5%75.0%
GET /post47.6%63.1%
GET /messages40.5%59.4%
GET /user51.0%48.6%

That work became the performance baseline for the new backend. Our compatibility requirement applied to latency as well as response shapes: Ayrshare 3 had to be at least as fast as the system it replaced, and that system had just become substantially faster.

Comparing the two weeks before the load balancer switch on 19 August with the two weeks after showed the same p95 latency on those endpoints.

The rebuild’s own return shows up differently. Since the migration finished, the MCP servers, link sessions for the Connect flow, the Automations engine and the WhatsApp Business beta have all shipped on the new code. Several would have been impractical on the previous backend.

The July work made API responses faster. Ayrshare 3 made it easier to ship new capabilities. For customers, that means new platforms and features sooner, and faster fixes when a social network changes its rules.

Completing the Migration and Keeping a Rollback Option

With the last background job moved on 27 August, the migration was done.

Migrations accumulate machinery. The real test of a team is whether the machinery gets removed afterward. The largest piece of ours, a few thousand lines built so that scheduled jobs could be moved one at a time, did its job and was deleted in early September once a single cloud command could do the same thing.

The previous backend is still deployed as a known-good copy that the load balancer can return to in seconds. Retiring it is the final, irreversible step, and we are not in a hurry to take it.

Build With the Ayrshare Social Media API

If you are building social publishing, analytics or messaging into your product or AI agents, explore the Ayrshare API documentation to see what you can build.

The 28-day trial on the Launch plan lets you try the API with your own traffic.

Questions about the rebuild itself are welcome too. Write to me at christos@ayrshare.com.

Written by: Christos Melas

Christos is a full-stack senior engineer and architect at Ayrshare. With 20+ years building enterprise distributed systems, he writes about social media APIs, architecture, and integrating with social platforms.

Christos Melas

Christos Melas

Engineering Manager

Christos is a full-stack senior engineer and architect at Ayrshare. With 20+ years building enterprise distributed systems, he writes about social media APIs, architecture, and integrating with social platforms.

Ask AI about this article

Opens your AI assistant in a new tab with this article preloaded.

FAQs About the Ayrshare API Backend Rebuild

No, the Ayrshare backend migration does not require changes to your integration. Routes, request bodies, responses and error codes are unchanged. The only differences are a few documented improvements to webhook delivery and rate limiting, covered in the webhooks and rate limiting docs.

No, Ayrshare still uses api.ayrshare.com, and there is no new API version to adopt. Ayrshare 3 is the new backend serving that existing API.

Accounts were moved behind the existing hostname with no change to their integration. The one customer-visible event was 42 gateway errors over two minutes when the load balancer was repointed.

Yes, the Ayrshare API is faster following separate performance work shipped in July 2026. That work reduced average and p95 latency on the most-used endpoints, including an 88.7% reduction in p95 latency for GET /profiles and 75% for GET /analytics. Ayrshare 3 maintained that improved performance baseline.

A parallel rebuild under a strict parity rule let us restructure everything at once while the existing backend kept serving traffic, and let us do it while that backend was healthy rather than under pressure.

We required the new backend to match the existing API’s routes, request bodies, response shapes, error codes, headers and database behavior. We compared every REST handler with its counterpart before routing customer traffic to the new system, then used feature flags to move accounts gradually.