Everything between the mess and the machine.
Four service lines. Sell one or the whole chain.
Migration — CMS and PDF to DITA
The hard part isn't converting markup. It's deciding what each piece of content means, so the conversion can be automated. That's where the effort goes, and it's why the pilot stage exists.
What gets in the way
Shortcodes with no DITA equivalent, page-builder markup wrapping every paragraph in three
nested divs, and content living in plugin tables rather than post_content. We
inventory shortcodes first and agree a disposition for each: convert, drop, or raise for an
author to rewrite.
What gets in the way
Intro-text and full-text splits that don't match any semantic boundary, and category trees used for navigation rather than meaning. We rebuild the hierarchy as a DITA map from the content itself, not from the menu structure.
What gets in the way
Storage-format macros, tables used for visual layout, and pages mixing reference material with meeting notes. We map the macros that carry meaning, strip the ones carrying only styling, and flag pages whose content type is genuinely ambiguous.
What gets in the way
Content fragmented across templates, renderings and personalisation rules. The upside: Sitecore's field structure is already semi-structured, so field values often convert directly into DITA conditional attributes — the fastest route to real profiling.
What gets in the way
PDFs carry layout, not structure. Recovery depends on how consistently styles were applied. We assess a sample before quoting, because a well-styled FrameMaker export and a scanned manual are the difference between automation and retyping.
Loose HTML, DOCX, Markdown
Drupal, Zendesk, SharePoint, FrameMaker, MadCap and in-house systems. If it exports to XHTML, XML, DOCX or Markdown, metR can be given a rule pack for it.
metR does the conversion. We do the judgement.
metR is Metapercept's structured content migration platform — rule-driven conversion from unstructured sources into DITA, custom DITA-OT plugin builds, and continuous publishing. As a partner company we configure it against your client's information model rather than running it on defaults.
Rule-pack conversion
Mapping rules are written once per source system and information model, then applied across the library. Consistency comes from the ruleset, not from reviewer discipline.
Reuse detection
Near-duplicate passages are identified during conversion and lifted into a shared library with conref pointers, instead of being carried forward as forty copies.
Custom DITA-OT plugins
Output styling, layout and compliance formatting built as plugins, so PDF and HTML5 match brand and regulatory requirements straight out of the build.
CI/CD publishing
Authoring, validation and publishing run as an automated pipeline against the repository, so documentation ships on the product's cadence.
Information architecture
A migration without a model just moves the mess into XML. This stage decides whether the next five years are easier than the last five.
| Deliverable | The decision it locks down | Why it matters later |
|---|---|---|
| Topic type model | Which content is a task, concept, reference, or a specialisation | Determines what validates automatically and what a human must read |
| Metadata dictionary | Controlled vocabularies for product, release, audience, lifecycle, owner | Drives faceted search, conditional publishing and AI retrieval quality |
| Taxonomy & subject scheme | The concept hierarchy those metadata values come from | Keeps tagging consistent as the library grows and authors change |
| Reuse strategy | What becomes a conref, a keyref, or a shared warehouse topic | Sets the ceiling on how much authoring and translation cost can be removed |
| Conditional attributes | Which axes you profile on — audience, product, platform, region | One source serves every variant instead of forking per market |
| Map architecture | How maps, submaps and keyscopes are organised | Determines whether two teams can work the library at once without collisions |
| Naming & ID conventions | File names, topic IDs, key names, folder structure | Unglamorous, and the most common cause of unmaintainable DITA |
Content operations
The migration is finite. Operations is what stops the library drifting back to the state it was in when you called us.
Authoring support
Structured authoring in oXygen or the client's CCMS, inside their model, to their style guide. Overflow capacity or a standing documentation team.
Review & approval workflow
Defined roles, review gates and sign-off records — configured in the CCMS, or as a Git pull-request workflow if DITA lives in a repository.
Governance & audit
Ownership metadata, review-due dates, change history, and a standing report on content past its review interval. Compliance stops being a fire drill.
Localisation handoff
Translation packages prepared from DITA source with memory alignment, so vendors quote on changed segments. Includes Arabic and right-to-left output handling.
Quality assurance
Automated validation on every commit — DTD conformance, Schematron business rules, link integrity, terminology, metadata completeness — plus scheduled editorial review.
CCMS support
Tool-agnostic. Heretto, RWS Tridion Docs, MadCap IXIA, Adobe AEM Guides, or oXygen with Git. We'll advise on selection if the client hasn't chosen.
Publishing and DITA-OT
Structured content only pays off at the point of delivery. It's also the part most often underestimated at scoping time, so price it in from the start.
Brand-accurate print and screen PDF via custom plugins, including regulatory layout requirements and controlled document furniture.
HTML5 & portals
Responsive help sites with faceted search, version switching and topic-level deep linking into support tooling.
In-product & headless
Content delivered by API into applications, embedded help panels, devices and chat interfaces — from the same source.
Build automation
Publishing wired into GitHub Actions, GitLab CI, Azure DevOps or Jenkins, triggered by merge or release tag.
Three ways in.
Wholesale pricing is quoted per engagement after the audit; partners set their own resale price. Durations assume a mid-sized library — the audit replaces them with real numbers.
Content audit
Fixed fee, fixed scope. Inventory, duplication analysis, risk register and a costed migration plan. The deliverable is yours whether or not anyone proceeds.
Typically 2–3 weeks
Pilot migration
One product line converted end to end, with a tuned metR rule pack and a working publishing output. Accuracy measured, not asserted.
Typically 4–6 weeks
Full migration & managed ops
Complete library conversion, publishing pipeline, training and handover — followed by a managed content operations retainer if the client wants it run for them.
Scoped from pilot results
Send a sample. We'll tell you what converts.
A CMS export, a handful of URLs, or one representative PDF is enough for a first read on feasibility and effort.