Technical Debt at Scale: Platform Migrations and Deprecations
Lesson 68: Technical Debt at Scale: Platform Migrations and Deprecations
Lesson 68: Technical Debt at Scale: Platform Migrations and Deprecations
Lesson 67 closed with a company that, having redesigned its trust and safety enforcement pipeline, faced a significant technical rework — precisely the kind of large-scale migration this lesson now addresses directly. Module 4 introduced the Technical Debt Quadrant as a way of categorizing debt at the level of an individual team's codebase. This lesson scales that concern up by an order of magnitude: what happens when the thing that needs to change is not a single team's internal implementation, but a Layer 2 Developer Surface, per the Leverage Stack from Lesson 61, that dozens or hundreds of external parties depend on, per the Promise Tiers model from Lesson 62.
At platform scale, technical debt is not merely an engineering inconvenience to be paid down at a team's own convenience — it can become a binding constraint on the platform's ability to evolve at all, because every dependent integration built by an external party is, in effect, a small claim on the platform's future flexibility. A platform that accumulates enough of these claims, without a disciplined process for eventually retiring old commitments, can find itself unable to make even clearly beneficial changes, trapped by the sheer volume and diversity of things depending on the old behavior.
This lesson introduces the Sunset Runway, this lesson's core mental model, to give you a structured way to plan and execute large-scale migrations and deprecations without either freezing a platform in place indefinitely or breaking trust with the ecosystem depending on it — directly extending the discipline this curriculum began building in Lesson 62's Promise Tiers.
Learning Objectives
- 1
Explain why platform-scale technical debt differs from team-level technical debt in kind, not just degree.
- 2
Apply the Sunset Runway model to plan a large-scale migration or deprecation with an appropriate timeline.
- 3
Identify the risks of an incomplete dependency inventory before a migration begins.
- 4
Describe the purpose of a dual-run period in a platform migration.
- 5
Evaluate a proposed migration plan for whether its timeline and communication match the scale of ecosystem dependency involved.
This lesson assumes the Promise Tiers model, semantic versioning, and deprecation policy concepts from Lesson 62, and a general familiarity with the Technical Debt Quadrant introduced in Module 4 (Lessons 31–40), since this lesson extends that team-level framework to the platform scale, where dependents are external parties rather than internal colleagues.
Why Platform-Scale Technical Debt Differs in Kind
Why Platform-Scale Technical Debt Differs in Kind
Team-level technical debt, as covered in Module 4's Technical Debt Quadrant, is primarily a matter of internal trade-offs: a team accepts some debt to move faster, and pays it down later, largely on its own schedule, since the people affected by the debt and the people who incurred it are often the same people, or at least the same organization. Platform-scale technical debt is structurally different, because the parties depending on the old behavior are frequently external, numerous, and invisible to the platform team in aggregate. A single internal team's debt is a private matter; a platform's accumulated commitments are, in effect, public promises made to an entire ecosystem, and unwinding them requires coordinating changes across parties the platform team does not control and may not even be able to fully enumerate.
This is precisely why Lesson 62 introduced the Promise Tiers model and formal deprecation policy: without those mechanisms, a platform has no legitimate way to ever retire an old commitment, and technical debt at this scale becomes a one-way ratchet, accumulating indefinitely because removing anything risks breaking someone, somewhere, who was never accounted for.
The Sunset Runway
The Sunset Runway
This lesson introduces the Sunset Runway, a four-phase model for retiring a platform capability that external parties depend on:
Phase 1, Dependency Inventory, requires building the most complete possible picture of who actually depends on the capability being retired, and how — a step that is frequently skipped or under-resourced, with damaging consequences illustrated in the Case Study below. Phase 2, Announcement and Dual-Run, begins the deprecation clock (per the Promise Tier the capability belongs to) while keeping both the old and new paths simultaneously functional, giving dependents time to migrate without a hard cutover forcing everyone to move at once. Phase 3, Active Migration Support, is where the platform team takes on real responsibility for helping dependents actually complete their migration — providing tooling, direct outreach to known high-impact dependents, and clear deadline communication — rather than simply announcing the change and waiting passively. Phase 4, Decommission, removes the old path only after the notice period has elapsed and, ideally, after monitoring shows migration completion has reached an acceptable threshold, rather than by calendar date alone regardless of actual migration progress.
The Sunset Runway's core discipline is recognizing that the length of an appropriate runway is not a fixed constant — it should scale with the number, diversity, and criticality of dependents uncovered in Phase 1, meaning the very first phase directly determines whether every subsequent phase is even planned realistically.
The Danger of an Incomplete Dependency Inventory
The Danger of an Incomplete Dependency Inventory
A dependency inventory built only from what the platform team can easily observe (registered API keys, documented integrations) will systematically miss dependents using undocumented workarounds, third-party tools built on top of official integrations, or long-tail, low-volume-but-critical usage that doesn't show up prominently in aggregate metrics. An incomplete inventory doesn't just risk missing a few edge cases — it risks systematically under-estimating the true scope of a migration, leading to a runway that looks generous on paper but is, in practice, far too short for the dependents the platform team never accounted for in the first place.
Why a Dual-Run Period Matters
Why a Dual-Run Period Matters
A dual-run period — during which both the old and new capability remain simultaneously functional — exists specifically to decouple the platform's readiness to retire something from each individual dependent's readiness to migrate away from it. Without a dual-run period, a migration becomes a synchronized, all-or-nothing cutover, forcing every dependent, regardless of their own internal priorities, resourcing, or release cycles, to migrate on the platform's timeline rather than a timeline that accounts for their own constraints — a mismatch that is especially costly for the platform's most resource-constrained and often most loyal long-tail dependents.
Common Mistakes to Avoid
Skipping or under-resourcing the dependency inventory phase
Assuming that documented, registered integrations represent the full universe of dependents nearly always understates the true scope of who will be affected by a migration.
Setting a migration deadline based on calendar convenience rather than dependent migration progress
Decommissioning on a fixed date regardless of actual migration completion rates converts a planned, orderly transition into a forced, disruptive cutover for anyone who hasn't finished migrating.
Treating announcement as equivalent to active migration support
Simply publishing a deprecation notice and expecting dependents to act on their own initiative, without direct outreach or migration tooling, tends to leave a long tail of dependents who never see or act on the announcement until the deadline arrives.
Underestimating the runway needed for the most resource-constrained dependents
A migration timeline appropriate for large, well-resourced partners may be far too aggressive for smaller developers or long-tail integrations with less engineering capacity to respond quickly.
Failing to monitor migration progress during the dual-run period
Without tracking how many dependents have actually migrated as the deadline approaches, a platform team has no way to know whether a planned Phase 4 decommission is actually safe to execute on schedule.
Ready to test your product judgment?
Take the interactive practice quiz for Lesson 68 and build your skill radar dashboard.