Stop Treating Headless CMS Migration Like a Copy-Paste Job: A Strategic Guide to Schema Translation
A headless CMS migration is never just a simple content transfer. In reality, it is a complex schema translation between two fundamentally incompatible data models. Treating it like a routine copy-paste job is precisely why so many migrations stall mid-flight or, worse, get quietly rebuilt from scratch six months after launch. The failures aren't usually mysterious; they stem from predictable issues like field-type mismatches, flattened relationships that need remodeling, locale structures that don't align, and asset URLs hardcoded into rich text.
At Orbitcore, we’ve seen that successful migrations hinge on sequencing. Each potential break must be caught during a rigorous audit phase, not in the heat of production. This guide breaks down how to navigate these technical hurdles to ensure your content model survives the move intact.
The Reality of Schema Translation
Schema translation is the core operation that defines whether a migration succeeds or fails. The friction usually starts with a field-type mismatch. For instance, a permissive WordPress rich-text field that happily accepts nested HTML and inline shortcodes has no direct equivalent in a strict environment like Contentful, Strapi, or Storyblok. Nothing flags this discrepancy until your validation script rejects the import and the project grinds to a halt.
Our engineering teams always audit legacy CMS exports for field-type mismatches, orphaned shortcodes, and reference-graph depth before any cutover. The pattern is consistent: content that looks perfectly clean in a CSV or JSON export often fails the moment it hits a structured, typed schema. Before you write a single line of migration code, you must diff every field type in your source against the target platform's library. Deeply nested reference chains aren't just things to map; they are structures that often need entirely new models.
Why the 'Blob' Approach Fails
A traditional WordPress post is essentially a 'blob.' It contains a title, a body of permissive rich text, some meta fields, and perhaps some shortcodes that only render correctly within the WordPress ecosystem. Modern headless systems like Storyblok or Contentful expect the opposite: structured content. They demand typed fields, validated formats, and explicit reference fields that link entries together rather than IDs buried inside a text string.
Moving between these worlds means re-expressing every content type as a stricter, more explicit schema. Contentful, for example, is very direct about this: a field’s type and validation rules are fixed. If the content doesn't conform, the API rejects it. You cannot coerce it. The unit of work here is the 'bounded content type.' You map one type at a time, resolving mismatches and orphaned references before moving to the next.
The Pre-Migration Structural Inventory
A pre-migration audit isn't a content audit; it’s a structural inventory. You need to identify what your source model will not survive. This audit should happen before you even sign off on a target platform, as the results will determine the true scope of your project.
Field-type mismatches are the primary culprit. A WordPress rich text field storing nested HTML and inline styles will fail when pointed at a Contentful rich text field. Why? Because Contentful stores content as a structured JSON document validated against specific node types. Anything outside that schema—like a <span> with inline CSS or a random [gallery] shortcode—either gets dropped or blocks the save entirely. You’ll often see this surface as a 422 error or a schema validation exception.
Remodeling Relationships and Reference Graphs
This is where original scope estimates usually fall apart. In a headless model, a shortcode-based gallery or a 'related posts' block isn't just text; it's a graph of linked entries. Rebuilding this requires a decision: which target content type does each shortcode map to?
Complexity grows with the depth of the reference graph. An article pulling in an author bio, which pulls in a social profile, which links to a product card, can easily end up five levels deep. While Contentful’s API can resolve references up to 10 levels deep, a healthy model rarely needs that much. Strapi, meanwhile, forces you to define cardinality (one-to-one, one-to-many) upfront—a decision legacy systems often never required you to make. You must count the average depth per content type during your audit and treat anything exceeding four levels as a candidate for its own isolated migration.
The Locale Logic: Per-Entry vs. Per-Field
Choosing the wrong localization model can force a total structural rebuild later. Strapi generally uses a locale-per-entry model, where each language gets its own entry linked back to a default. Reference fields then point to a specific language version.
Contentful often takes the opposite route: locale-per-field. One entry holds French and English values side-by-side, and references point to the entry itself, resolved at the time of delivery. If your source CMS stores translations as separate posts, mapping them to a locale-per-field model requires a migration script that merges sibling records before the translation even starts. Guessing wrong here is a mistake that might not surface for months, but when it does, it’s painful.
Saving Your SEO: Assets and Redirects
Asset URL migration is frequently treated as a formality, but it’s actually where SEO goes to die. WordPress bakes absolute URLs directly into the content. When you migrate to a headless provider, those assets are rehosted on new domains (like images.ctfassets.net or a.storyblok.com). If you don't update those links, your frontend renders broken images and search engines start de-indexing your pages.
You need two fixes: first, a script to rewrite embedded URLs based on an old-to-new asset ID map. Second, a robust 301 redirect strategy. A 301 redirect is the only way to tell Google that a URL change is permanent. Skip this, and you might keep your content but lose all your hard-earned traffic.
Fiber network designs you can actually rely on.
We handle the heavy lifting. From local surveys in Java & Medan to detailed FTTH grid designs, we make sure your network makes sense.
Bridging the Editorial Workflow Gap
There is an 'invisible' workflow that editors use daily which often disappears during a headless migration. In traditional CMS platforms, drafting, previewing, and publishing happen on one screen. In a headless setup, these are decoupled. Content lives in an API, and the 'preview' is a separate frontend build.
Before migrating, list every action an editor takes. Can they preview unpublished entries? Is there a scheduled publish feature? Contentful, Strapi, and Storyblok all handle draft states differently. You must rebuild the editorial process to match the new platform's logic, not just move the data fields.
The Power of the Parallel Run
Never migrate everything at once. The safest path is to move one bounded content type first—ideally something simple like an FAQ or a basic landing page with low traffic. Migrate it end-to-end: content, assets, and the editorial workflow.
Then, run both systems in parallel. Let the new CMS serve production traffic for a specific section of the site for two to four weeks. This 'parallel run' is the only way to catch real-world issues like broken reference resolutions or locale fallback quirks that testing might miss. Only once the new system has proven stable—with zero unresolved references and a successful editorial cycle—should you cut over the rest of the site.
Final Thoughts
The most expensive parts of a headless migration are the decisions made before the first script is written. Understanding how fields map, how references flatten, and how locales are structured will determine if your project takes two weeks or six months. If you’re at the stage where your target platform is chosen but your content isn't mapped, that mapping is the most valuable work you can do right now. Don't wait for the export to fail to start thinking about the schema.