The vendor's docs are a PDF from 2009. The mainframe extract is fixed-width text with no header row. The "API" is a SOAP endpoint that returns German error messages. Welcome to the integration work that actually pays FDE salaries — and the translation-layer pattern that keeps the mess from infecting everything you build.

Assumes: Posts 1 (REST), 2 (auth), and 7 (SDK design). Installs: pip install httpx. The SOAP server, the fixed-width file, and the mystery API are all built below — every byte shown is generated by code you can run.

Thursday, 10:15 AM. Maria: "Good news and bad news. Good news: the sanitation department wants their crew assignments in CityOps. Bad news: their system was built in 1998, the vendor went bankrupt in 2011, and the only documentation is a PDF that says 'see appendix B' — there is no appendix B."

You: "What does the system actually expose?"

Maria: "Three things. A SOAP endpoint for crew assignments. A nightly fixed-width file dump from the mainframe. And an internal web page that the last engineer scraped — nobody knows if it's an API or an accident."

Dev: "So we write three integrations?"

You: "No. We write one boundary. Three adapters behind it, one clean interface in front of it. The mess stays behind the boundary — forever."

The minimum concept

A translation layer (also called an anti-corruption layer, after the Domain-Driven Design pattern) is a boundary between your clean domain and someone else's mess. Everything on your side speaks your language — Crew, Assignment, ISO dates. Everything on their side stays their problem — SOAP envelopes, fixed-width columns, mystery JSON. The layer translates, validates, and quarantines. Nothing leaks through untranslated.

Legacy interfaceIts native shapeWhat your side sees
SOAP endpointXML envelopes, WSDL, German faultsget_assignments(date) → list[Assignment]
Mainframe extractFixed-width text, no headers, EBCDIC memoriesget_crews() → list[Crew]
Undocumented web pageHTML meant for browsersget_status(id) → str

The design principle: meet the legacy system where it is — and keep its mess behind the boundary. You don't fix the mainframe. You don't rewrite the SOAP service. You build a wall with a very clean door, and your code only ever walks through the door.

The translation-layer contract

  • One adapter module per legacy interface — never one file with three hacks
  • Adapters return domain objects, never raw legacy shapes
  • Every adapter validates what it parses; unparseable input is quarantined, not guessed at
  • The boundary is strict: legacy concepts (record types, SOAP faults, screen-scrape selectors) never appear in calling code
  • Be liberal in observation, strict at your boundary — log everything the legacy system does, accept only what validates

Build it, part 1: the SOAP adapter

First, the legacy SOAP service itself — a mock with authentic habits: XML envelopes, a WSDL nobody reads, and fault messages in German:

"""Mock legacy SOAP service: crew assignments, 1998 vintage."""
import threading
from http.server import BaseHTTPRequestHandler, HTTPServer

SOAP_OK = """<?xml version="1.0"?>
<soap:Envelope xmlns:soap="http://schemas.xmlsoap.org/soap/envelope/">
  <soap:Body>
    <GetAssignmentsResponse xmlns="http://sanitation.cityops.example/">
      <Assignment><Id>ASN-101</Id><CrewId>CRW-001</CrewId>
        <Route>R-201</Route><Date>2026-10-08</Date></Assignment>
      <Assignment><Id>ASN-102</Id><CrewId>CRW-002</CrewId>
        <Route>R-202</Route><Date>2026-10-08</Date></Assignment>
    </GetAssignmentsResponse>
  </soap:Body>
</soap:Envelope>"""

SOAP_FAULT = """<?xml version="1.0"?>
<soap:Envelope xmlns:soap="http://schemas.xmlsoap.org/soap/envelope/">
  <soap:Body>
    <soap:Fault>
      <faultcode>soap:Server</faultcode>
      <faultstring>Ungültige Anfrage: Datumsformat</faultstring>
    </soap:Fault>
  </soap:Body>
</soap:Envelope>"""

class SoapHandler(BaseHTTPRequestHandler):
    def do_POST(self):
        length = int(self.headers.get("Content-Length", 0))
        body = self.rfile.read(length).decode()
        # the legacy service validates the date format strictly: YYYYMMDD
        if "<Date>" in body and body.split("<Date>")[1].startswith("2026"):
            payload, code = SOAP_OK, 200
        else:
            payload, code = SOAP_FAULT, 500
        raw = payload.encode()
        self.send_response(code)
        self.send_header("Content-Type", "text/xml")
        self.send_header("Content-Length", str(len(raw)))
        self.end_headers()
        self.wfile.write(raw)

def run_soap(port=8941):
    srv = HTTPServer(("127.0.0.1", port), SoapHandler)
    threading.Thread(target=srv.serve_forever, daemon=True).start()
    return srv

Now the adapter. Note what it does that Post 7's SDK didn't need to: it speaks the legacy protocol's date dialect (YYYYMMDD), parses XML with namespaces, and translates the German fault into a typed exception with a boundary error that names the translation:

"""soap_adapter.py — one adapter for the 1998 SOAP service."""
import httpx
import xml.etree.ElementTree as ET
from dataclasses import dataclass

@dataclass
class Assignment:
    id: str
    crew_id: str
    route: str
    date: str   # ISO YYYY-MM-DD on our side of the boundary, always

class LegacyError(Exception):
    def __init__(self, message, *, raw=None):
        super().__init__(message)
        self.raw = raw   # quarantined for debugging, never shown to callers

NS = {"s": "http://schemas.xmlsoap.org/soap/envelope/",
      "t": "http://sanitation.cityops.example/"}

def required_text(elem, path):
    """ElementTree returns None for missing nodes — turn that into a
    boundary error that names the missing field, not a None downstream."""
    child = elem.find(path, NS)
    if child is None or child.text is None:
        raise LegacyError(
            f"SOAP response missing required field '{path}'; "
            f"quarantined {ET.tostring(elem)[:120]!r}")
    return child.text

def get_assignments(date_iso, base_url="http://127.0.0.1:8941"):
    # the legacy dialect: YYYYMMDD. The boundary translates both ways.
    legacy_date = date_iso.replace("-", "")
    envelope = f"""<?xml version="1.0"?>
<soap:Envelope xmlns:soap="http://schemas.xmlsoap.org/soap/envelope/">
  <soap:Body>
    <GetAssignments xmlns="http://sanitation.cityops.example/">
      <Date>{legacy_date}</Date>
    </GetAssignments>
  </soap:Body>
</soap:Envelope>"""
    resp = httpx.post(base_url, content=envelope,
                      headers={"Content-Type": "text/xml",
                               "SOAPAction": "GetAssignments"},
                      timeout=10)
    root = ET.fromstring(resp.text)
    fault = root.find(".//s:Fault", NS)
    if fault is not None:
        detail = fault.findtext("faultstring")
        raise LegacyError(
            f"legacy SOAP fault: {detail} "
            f"(boundary: GetAssignments for {date_iso})",
            raw=resp.text)
    out = []
    for node in root.findall(".//t:Assignment", NS):
        out.append(Assignment(
            id=required_text(node, "t:Id"),
            crew_id=required_text(node, "t:CrewId"),
            route=required_text(node, "t:Route"),
            date=date_iso))   # our canonical form, not theirs
    return out
asns = get_assignments("2026-10-08")
print([(a.id, a.crew_id, a.route, a.date) for a in asns])
# [('ASN-101', 'CRW-001', 'R-201', '2026-10-08'),
#  ('ASN-102', 'CRW-002', 'R-202', '2026-10-08')]

try:
    get_assignments("not-a-date")
except LegacyError as e:
    print("LegacyError:", e)
# LegacyError: legacy SOAP fault: Ungültige Anfrage: Datumsformat
#   (boundary: GetAssignments for not-a-date)

The German fault arrives, and the caller gets an English exception naming the boundary, the operation, and the input — with the raw XML quarantined on e.raw for the ticket, not splashed across the logs. required_text is the small function that does the most work: a missing XML node becomes a loud boundary error instead of a None that poisons downstream code.

Build it, part 2: the fixed-width mainframe extract

The nightly dump. No headers, no delimiters — every field lives at a fixed column offset, exactly as the COBOL program wrote it in 1998:

HD20261008SANITATION   EXTRACT
CRW-001   Alpha     north day
CRW-002   Bravo     south day
CRW-003   Charlie   north night
DT20261008SANITATION   EXTRACT
TR0000003

Read it like the mainframe does — by position, not by parsing. The spec (reconstructed from the PDF's appendix that does exist, plus one phone call with a retired operator):

# positions are 0-indexed, end-exclusive — measured, not guessed
SPEC = {
    "HD": [("record_type", 0, 2), ("date", 2, 10), ("system", 10, 31)],
    "DT": [("record_type", 0, 2), ("crew_id", 2, 12), ("name", 12, 22),
           ("zone", 22, 28), ("shift", 28, 33)],
    "TR": [("record_type", 0, 2), ("count", 2, 9)],
}
KNOWN_ZONES = {"north", "south", "east", "west"}
KNOWN_SHIFTS = {"day", "night"}
"""fixedwidth_adapter.py — one adapter for the mainframe extract."""
from dataclasses import dataclass

@dataclass
class Crew:
    id: str
    name: str
    zone: str
    shift: str

class Quarantine(Exception):
    """Raised with the offending line attached — the nightly job catches
    this per-line, logs it, and keeps going."""
    def __init__(self, message, *, line=None):
        super().__init__(message)
        self.line = line

def parse_line(line, spec_name):
    # the 31-character contract: short lines are corrupt, not "flexible"
    if len(line.rstrip("\n")) != 31:
        raise Quarantine(
            f"{spec_name}: expected exactly 31 chars, "
            f"got {len(line.rstrip())}", line=line)
    line = line.rstrip("\n")
    if line[0:2] not in ("HD", "DT", "TR"):
        raise Quarantine(f"unknown record type {line[0:2]!r}", line=line)
    fields = {}
    for name, start, end in SPEC[spec_name]:
        fields[name] = line[start:end].strip()
    return fields

def parse_extract(text):
    lines = [ln for ln in text.splitlines() if ln.strip()]
    header = parse_line(lines[0], "HD")
    crews, errors = [], []
    for ln in lines[1:]:
        rt = ln[0:2]
        if rt == "TR":
            trailer = parse_line(ln, "TR")
            continue
        try:
            f = parse_line(ln, "DT")
            # presence-based validation: a field is either known-good
            # or the line is quarantined — no silent .get() chains
            if f["zone"] not in KNOWN_ZONES:
                raise Quarantine(f"unknown zone {f['zone']!r}", line=ln)
            if f["shift"] not in KNOWN_SHIFTS:
                raise Quarantine(f"unknown shift {f['shift']!r}", line=ln)
            crews.append(Crew(id=f["crew_id"], name=f["name"],
                              zone=f["zone"], shift=f["shift"]))
        except Quarantine as q:
            errors.append(q)
    # the trailer is a checksum on the file's honesty: reconcile it
    # against DT records actually parsed, not lines counted
    if int(trailer["count"]) != len(crews) + len(errors):
        errors.append(Quarantine(
            f"trailer count {trailer['count']} != parsed {len(crews)} "
            f"+ quarantined {len(errors)}"))
    return {"date": header["date"], "crews": crews, "errors": errors}
result = parse_extract(EXTRACT_TEXT)
print("date:", result["date"], "| crews:", len(result["crews"]),
      "| errors:", len(result["errors"]))
# date: 20261008 | crews: 3 | errors: 0

Three crews parsed, zero quarantined, trailer reconciled. Now watch the adapter earn its keep — a corrupt line and a lying trailer:

bad = EXTRACT_TEXT + "CRW-999   Zulu      up    day\n"   # 28 chars, bad zone
result = parse_extract(bad)
print("crews:", len(result["crews"]), "| errors:", len(result["errors"]))
for e in result["errors"]:
    print("quarantined:", e)
# crews: 3 | errors: 2
# quarantined: DT: expected exactly 31 chars, got 28
# quarantined: trailer count 0000003 != parsed 3 + quarantined 1

The short line is quarantined with its exact length; the trailer mismatch is caught because the count reconciles against records actually parsed — and the nightly job logs both and keeps going. The pipeline doesn't die because one line is corrupt; the corrupt line doesn't silently become a crew named None.

Build it, part 3: the undocumented "API"

The third interface: an internal status page that returns JSON if you ask nicely. No docs, no versioning, no guarantees. The adapter treats it like what it is — a screen scrape with a seatbelt:

"""mystery_adapter.py — one adapter for the undocumented status page."""
import httpx

def get_status(crew_id, base_url="http://127.0.0.1:8942"):
    resp = httpx.get(f"{base_url}/status/{crew_id}", timeout=10)
    resp.raise_for_status()
    data = resp.json()
    # presence-based key normalization: the page has renamed keys twice
    # this year. Accept the known spellings; reject everything else loudly.
    for key in ("status", "state", "current_status"):
        if key in data:
            return data[key]
    raise LegacyError(
        f"status page returned no recognized status key; "
        f"keys seen: {sorted(data.keys())}", raw=resp.text)

Presence-based, not .get()-chained: data.get("status") or data.get("state") returns None for a payload with neither key, and None becomes "crew status: None" in someone's dashboard. The in check raises instead — because a missing status is information (the page changed again), not an empty value.

The orchestration: one clean door

Three adapters, one boundary. The orchestration module is the only thing calling code imports — and notice what it never mentions: SOAP, columns, selectors, record types:

"""sanitation_boundary.py — the one clean door."""
from soap_adapter import get_assignments, LegacyError
from fixedwidth_adapter import parse_extract, Quarantine
from mystery_adapter import get_status

def daily_sync(extract_text, date_iso):
    """Pull everything behind the boundary. Returns domain objects and
    a quarantine report — never raw legacy shapes."""
    report = {"assignments": [], "crews": [], "quarantined": []}
    try:
        report["assignments"] = get_assignments(date_iso)
    except LegacyError as e:
        report["quarantined"].append(("soap", str(e), e.raw))
    parsed = parse_extract(extract_text)
    report["crews"] = parsed["crews"]
    report["quarantined"].extend(("fixed-width", str(e), e.line)
                                 for e in parsed["errors"])
    return report
report = daily_sync(EXTRACT_TEXT, "2026-10-08")
print("assignments:", len(report["assignments"]),
      "| crews:", len(report["crews"]),
      "| quarantined:", len(report["quarantined"]))
# assignments: 2 | crews: 3 | quarantined: 0

Two assignments from 1998's SOAP service, three crews from the mainframe, zero quarantined — all as domain objects with ISO dates. The calling code knows nothing about envelopes, columns, or German. That's the boundary doing its job.

Break it: the legacy system moves under you

Legacy systems change without telling you — that's what makes them legacy. Three real failure modes, reproduced:

1. The SOAP service renames a field. <CrewId> becomes <Crew_Id> overnight. required_text raises LegacyError: SOAP response missing required field 't:CrewId' — naming the field, with the raw XML quarantined. The old code would have emitted Assignment(crew_id=None) and the None would have ridden into CityOps. Strict at the boundary beats clever downstream.

2. The mainframe adds a column. Lines grow from 31 to 35 characters. Every line fails the length contract — loudly, all at once, on the first nightly run after the change. That's the right failure: a length change means the spec changed, and silently parsing 35-char lines against 31-char offsets would misalign every field. The quarantine report becomes the spec-change detector.

3. The status page renames its key a third time. get_status raises LegacyError listing the keys actually seen. You add the new spelling to the accepted list — one line, in one adapter — and the calling code never notices.

Legacy integration rules

  • One adapter module per legacy interface — the mess never mixes
  • Validate at parse time; quarantine per-record, never abort the batch
  • Reconcile control records (trailers, counts) against what you actually parsed
  • Strict at the boundary: missing required data raises, never emits None
  • Don't make your integration depend on the legacy system changing — adapt at the boundary instead
  • Raw archives live under security/retention governance — quarantined payloads are evidence, and evidence has a retention policy

Productionize: living with the system

Monitor the boundary, not just the data. Two different alerts: schema alerts (a new record type, a renamed SOAP field, a line-length change — the contract moved) versus data-distribution monitoring (quarantine rate spiking from 0.1% to 5% — the data got sicker). The first pages the integration owner; the second pages the data team. Conflating them means the wrong person gets woken up.

Archive everything raw. Every SOAP response, every extract file, every status page body — stored raw, under the retention policy Lisa's team sets, before translation. When the 2 AM question is "what did the legacy system actually send?", the answer is a file, not a memory. Alerts link to quarantined inputs by reference — they never attach raw payloads, which may carry PII the alert channel isn't cleared for.

Version your adapters. The SOAP adapter pins the WSDL version it was built against; the fixed-width adapter pins the 31-char spec revision. When the legacy side changes, you change the adapter version — and the quarantine report tells you exactly when that day arrives.

Never "just" fix the legacy system. The temptation, every time: "while we're here, let's fix the mainframe extract." No. Your integration must not depend on the legacy system changing — the team that owns it has a 3-year backlog and no test environment. Adapt at the boundary; the boundary is yours.

Explain it to the customer

You, Thursday: "The sanitation department's three systems now feed CityOps through one boundary. Your team calls daily_sync() and gets crews and assignments as clean objects — no XML, no column offsets, no German. When their systems change — and they will — the change is caught at the boundary, quarantined with the evidence attached, and fixed in one adapter module. The mess stays behind the wall."

Maria: "And when they finally replace the mainframe?"

You: "We delete one adapter. The calling code doesn't change — that's what the boundary was for."

Useful later

  • Change-data-capture — when the nightly extract becomes a stream (Post 3's polite-client lesson, reversed)
  • Schema registries — when you own both sides of a boundary
  • Contract testing — pin the legacy behavior you depend on, the way Post 7's mock pinned the vendor
  • Gradual strangulation — replace legacy interfaces one adapter at a time

Myth: "Legacy integration is low-skill glue work."

Reality: It's the highest-leverage boundary design you'll do. The translation layer is where "the business runs on this" meets "we can actually change things" — and the engineer who builds that wall well becomes the person everyone trusts with the next migration. Glue work is what happens without the boundary.

Post 7's principle was hide the wire, not the error. This post's: meet the legacy system where it is — and keep its mess behind the boundary.

Field check

  1. The mainframe team announces 35-character lines starting next month. Walk through exactly what breaks, what the quarantine report shows, and what you change.
  2. A SOAP fault arrives with an empty <faultstring>. What does the caller see, and where would you improve the adapter?
  3. The status page starts returning HTTP 200 with an HTML login page (session expired). Which layer catches it, and what should it do?
  4. Someone proposes merging the three adapters into one file "since they're all sanitation." What's the real cost?
What good answers look like

1. Every line fails the 31-char contract on the first run — the quarantine report fills with length errors, which is the spec-change detector firing correctly. Nothing misaligns silently because parsing never proceeds on wrong-length lines. You change: the SPEC offsets, the length constant, the pinned spec revision — all in fixedwidth_adapter.py only. The trailer record (TR) also grows; its offsets change too. Calling code doesn't change. 2. The caller sees LegacyError: legacy SOAP fault: (boundary: GetAssignments for ...) — technically correct, practically useless. Improve the adapter: when faultstring is empty, fall back to faultcode and the raw envelope's first 200 chars in the message, so the exception still tells the debugger something. 3. The mystery adapter catches it: resp.json() raises (HTML isn't JSON) — currently an unhandled json.JSONDecodeError, which is the gap to fix: catch it, raise LegacyError("status page returned non-JSON; session may have expired") with the raw body quarantined. It should alert (schema alert — the contract moved) and the adapter should support re-authentication or cookie refresh. 4. The mess mixes: a SOAP change forces you to re-read fixed-width code; the quarantine report loses per-interface attribution; testing one adapter requires loading all three; and when the mainframe is finally replaced, you can't delete one file — you perform surgery. One adapter per interface is what makes "delete one adapter" possible.