All work

Case study · 202501 / 06

UXCam

Moved a mobile analytics platform from a PHP monolith to Python microservices in six months, then built its billing system and a 100K-a-day event queue.

Role
Senior Backend Engineer
Year
2025
Stack
  • Python
  • Django
  • PostgreSQL
  • Redis
  • AWS ECS
  • Docker
  • OAuth2/SSO
  • Event queues
UXCam marketing site showing its mobile app analytics product, with session recordings and analytics dashboards for iOS and Android apps

Problem

UXCam records and analyses user sessions in iOS and Android apps in real time. When I joined in September 2023, the backend was a legacy PHP monolith. It was slow to change and expensive to run, and customer data was spread across several business tools with no single view of a customer.

On top of that, the authentication service was failing about 15% of logins, and the company needed a proper subscription system to bill customers.

Every one of those problems touched revenue. Slow releases meant slow fixes. Scattered customer data meant sales, support and success teams each saw a different picture. And a login that fails one time in seven is the fastest way to lose a customer’s trust.

My role

I’m a Senior Backend Engineer. I led the migration from the PHP monolith to Python microservices, built the event queue that feeds our business tools, designed the subscription system, and look after the core authentication service (OAuth2 and SSO).

Leading the migration meant owning the plan as much as the code: deciding the order services came out of the monolith, setting the patterns each new service followed, and making sure the product kept working for customers the whole way through.

Architecture

A legacy PHP monolith was migrated to Python microservices running on Django, Docker and AWS ECS. The services include an OAuth2 and SSO auth service, a subscriptions service for billing and dunning, and an event queue handling over 100,000 events a day. Auth and subscriptions use PostgreSQL and Redis; the event queue syncs to HubSpot, Mixpanel, Intercom and Planhat.
The monolith was replaced service by service; the event queue is what gives the business one view of each customer.
  • Python microservices in Django, packaged with Docker and running on AWS ECS, replaced the PHP monolith over six months.
  • Auth service handles OAuth2 and SSO for every login to the product.
  • Subscriptions covers recurring billing, invoicing, credits and dunning.
  • Event queue syncs more than 100,000 events a day to HubSpot, Mixpanel, Intercom and Planhat.
  • PostgreSQL and Redis hold service data.

The shape is deliberately boring. Each service owns its own data and has one job. Anything slow or unreliable, like a third-party API, sits behind the queue rather than in the path of a user request. Containers on ECS mean every service is built, deployed and scaled the same way, which is a big part of why releases got faster.

Key decisions & tradeoffs

Migrate in steps, not in one cutover. Moving a live analytics product off a monolith in six months meant moving it piece by piece rather than in one risky cutover. Each area came out on its own: build the Python service, run it alongside the PHP code it replaces, move traffic across, and only delete the old path once the new one has proven itself in production.

The rule I held to was that customers should never notice. A step that couldn’t be rolled back quickly wasn’t ready to ship. That made each step smaller and the whole migration feel slower at the start, but it meant no single change could take the product down, and the pace picked up as the patterns settled.

One queue for every downstream tool. Rather than each service calling HubSpot, Mixpanel, Intercom and Planhat directly, events go through a single queue. A slow or failing third-party API then delays a sync instead of breaking a user-facing request.

The design assumes things will fail. Delivery is at least once, so every sync has to be safe to run twice: events carry a stable ID, and handlers write in a way that makes a repeat a no-op. Rate limits and temporary errors are retried with backoff; anything that keeps failing is set aside for inspection instead of blocking the events behind it. At 100,000+ events a day, that is the difference between a queue you trust and one you babysit.

Find the race, don’t add retries. The login failures came from six separate concurrency bugs in the auth service. I tracked each one down and fixed it at the source rather than hiding the failures behind client-side retries.

Concurrency bugs don’t show up when you click through a login by hand. They show up when two requests touch the same session or token at the same moment. So the work was to make each failure reproducible under concurrent load first, then fix the underlying ordering or locking problem, and keep the reproduction as a test so it can’t quietly come back. Retries would have made the graph look better while leaving every one of those bugs in place.

Billing is correctness first. A subscription system that bills twice, or not at all, costs more trust than almost any other bug. I treated recurring billing, invoicing, credits and dunning as a set of explicit states with clear transitions, so every charge, credit and retry can be traced back to why it happened.

Outcome & metrics

  • Migration finished in six months.
  • Deployment frequency tripled.
  • Infrastructure costs down 35%.
  • 100K+ events a day synced to four business tools, giving one view of every customer.
  • Login errors cut from 15% to below 0.5% after fixing six concurrency bugs.

Beyond the numbers, the company now runs on a platform it can change with confidence. Teams ship more often because each release is smaller and safer. Sales, support and success work from the same customer data. And billing runs as a proper system with recurring charges, invoices, credits and dunning built in, rather than a manual process.

What I’d do differently

Put tracing in before the first split. Once a request crosses several services, logs alone don’t tell you where it went wrong. Distributed tracing from day one would have made cross-service problems, including the auth races, much faster to see.

Write the contracts down earlier. Between services and the queue, the shape of each event is a contract. I’d define and version those schemas up front, so a change in one service can’t silently break a consumer downstream.

Load-test auth before it hurts. The six concurrency bugs were found because logins were failing. A routine concurrent load test on the auth service would have surfaced them before customers did.