Migrating users between two live Django apps sounds like it needs one big migration script and a maintenance window. It turns out it doesn't.

As a user on the Pybites platform, you never had to do anything: just log in as usual, and if you had a v1 account, your premium status, exercise progress, and token balances were automatically imported into the new platform. No scripts, no maintenance window, just a series of API calls triggered by the login signal.

Here's how it worked.

The Problem: Two Platforms, One User Base

October 2024 the original Pybites platform, codechalleng.es (v1), showed its age, the UX was no longer up to standards. My first intuition was to go full API and build a separate front-end, but it soon became clear this would add a lot of complexity.

This was a major Django code base with many business logic decisions and hundreds of users with history. Complex user profiles; from free to exercise token access based to full premium access.

So I built a new Django app (v2) with a clean slate, migrating the exercises and meta data, and using Tailwind CSS + HTMX for a much better front-end experience. However, I needed to migrate users from v1 to v2. Users data: premium status, completed/attempted bites, token balances (earned/unlocked). Ideally this transition would happen without asking them to do anything.

The constraint that shaped everything

What worked really well was deciding early on to use separate databases and treat them as separate apps. This meant I had to solve the migration problem, but it also meant I had a clean slate to design the v2 data model without worrying about legacy constraints. And users could try v2 before making the switch, which made the transition smoother.

The naive approach and why I rejected it

When you're dealing with a massive migration, the first instinct is to write a big batch script: export all users from v1, transform the data, import into v2. But this has some downsides:

  • The data can become stale between export and import, especially if users are active during the migration window
  • Timing issues: you have to choose a cutover date, and any users who log in after the export but before the import will have inconsistent data

Instead I chose to migrate lazily, on first login. Each user migrated themselves when they logged in, one at a time. The data was always fresh, there was no need for a maintenance window. Users who never came back simply never migrated, no harm done.

The "Lazy migration upon login" design

This is the high-level flow I came up with:

User logs in on V2
        ↓
Django user_logged_in signal fires
        ↓
Check: v1_migration_done == False?
        ↓
POST V1:/api/v2/token/   ← prove you're the same user
        ↓
JWT access token returned
        ↓
PUT V1:/api/v2/bite-user-data/  (Bearer token)
        ↓
V1 returns: {profile, bite_tokens, bite_saves}
V1 marks user inactive
        ↓
V2 imports data → marks v1_migration_done = True

The key part was this Django signal handler in v2's accounts/models.py:

@receiver(user_logged_in)
def social_account_logged_in_handler(request, user, **kwargs):
    if settings.V1_MIGRATION_DISABLED:
        return
    if user.profile.v1_migration_done:
        return

    social = user.socialaccount_set.first()
    provider = social.provider  # "github" or "google"
    uid = social.uid

    payload = {"provider": provider, "uid": uid}
    # ... POST to v1, get JWT, PUT for data, import

There are three early returns here. The first two are shown above; the third bails out if the user has no social accounts, which brings us to the email/password case below. I ended up introducing an environment variable to temporarily re-open the migration process after the grace period expired. The v1_migration_done flag is the most important one; it ensures that once a user has migrated, they won't try to do it again on every login.

Email/password users needed a different hook. The signal only fired after a successful login, but these users didn't exist in v2 yet, so they never got that far. Instead, I overrode form_invalid in a custom LoginView. When login fails and the user doesn't exist in v2 yet, instead of immediately showing an error, it tries v1 with the same credentials. If v1 confirms them, it creates the account and imports the data on the spot.

class CustomLoginView(LoginView):
    def form_invalid(self, form):
        if settings.V1_MIGRATION_DISABLED:
            return super().form_invalid(form)

        email = self.request.POST.get("login")
        password = self.request.POST.get("password")

        # user not in v2 yet, try v1 with the same credentials
        response = httpx.post(f"{settings.PYBITES_CC_API_URL}/token/", ...)

        if response.status_code == 200:
            user = User.objects.create_user(...)
            import_user_v1_data_into_v2(user, ...)
            login(self.request, user, ...)
            return super().form_valid(form)

        return super().form_invalid(form)

Both paths call the same V1 endpoint, which returns:

# V1: api_v2/views.py
{
    "profile": {
        "premium": bool,
        "newbie_access": bool
    },
    "bite_tokens": {
        "earned": int,
        "unlocked_bite_ids": [42, 77, 203, ...]
    },
    "bite_saves": [
        {"bite_id": int, "code": str, "ok": bool, "added": "2023-..."}
    ]
}
# Side effect: V1 user set inactive after successful data retrieval

Error Handling

I had to handle 3 failure modes gracefully, since this was happening on login and I didn't want to break the user experience:

  1. User not found on V1 (404): User never had a v1 account; mark migration done, move on, no noise
  2. Auth failure: Log to a new MigrationErrorLog table, email admin, show neutral message to user
  3. Data import exception: Wrapped in transaction.atomic(); rollback on any failure, log it

There was one silent failure that shows how the devil is always in the details. The Google provider name was different between v1 and v2 ("google-oauth2" vs "google"), causing every Google user to fail to match, which was a silent failure since the migration just wouldn't happen, but the user would never know why. The fix was simply to add a normalizer in the v1 API endpoint: if provider == "google": return "google-oauth2"


What made it work in practice

  • The idempotency flag (v1_migration_done) is the simplest, most important thing. Without it, every login retries the migration.
  • Lazy migration meant we could run both platforms in parallel for weeks, no hard cutover date
  • Using the user's own OAuth credential as the auth mechanism meant no separate credential store, no "import token" to manage
  • V1 side effect (inactivating the user) was intentional and semantic: PUT, not GET, because it changes state

Users migrated on first login with no manual steps, no downtime, no stale data, and almost no support tickets.

When we shut down the process, users would still reach out asking why their progress didn't migrate, but that was expected and was easily handled by toggling the V1_MIGRATION_DISABLED flag to allow exceptional late migrations (which also shows the power of the 12 Factor App's recommendation of separating config from code).


What's a challenging migration you've had to do, and how did you approach it? Reach out and let me know.