Home›Read›Tactics desk›Parsing legacy PDFs: How Flipside optimized RAG for complex schematics
Tactics·Oct 4, 2026

Parsing legacy PDFs: How Flipside optimized RAG for complex schematics

When standard OCR turned dense pinball repair tables into unreadable text soup, Flipside rebuilt its retrieval pipeline around layout-aware parsing, platform-based document grouping, and isolated…

When standard OCR turned dense pinball repair tables into unreadable text soup, Flipside rebuilt its retrieval pipeline around layout-aware parsing, platform-based document grouping, and isolated synonym mapping.

Anthony, founder of Australian pinball marketplace Flipside, scaled a free AI repair assistant to 2,300 service manuals and 141,974 document chunks. The tool, hosted at joinflipside.com.au/repair, helps users diagnose complex mechanical and electrical faults across roughly 1,000 machines. Building a reliable retrieval-augmented generation (RAG) pipeline for legacy documents, however, exposed the limits of off-the-shelf vector search. Standard text extraction routinely failed on dense schematics, while naive keyword matching missed critical domain-specific jargon.

To make the tool functional, Anthony had to re-engineer the ingestion pipeline, restructure document metadata around hardware platforms rather than manufacturers, and isolate synonym expansion from full-text search. The resulting playbook provides a concrete blueprint for any developer building RAG applications on top of legacy, table-heavy technical documentation.

Layout-aware table parsing

Technical repair manuals from the 1990s rely heavily on dense tables to map switch matrices, coil tables, and fuse charts. Standard optical character recognition (OCR) engines extract these tables as continuous strings of text, destroying the structural relationships between rows and columns. To preserve this spatial context, Flipside routes every PDF through Azure Document Intelligence using the prebuilt-layout model.

This layout model extracts tables as structured cells, which are then written out as markdown format within each document chunk. While standard plain-text OCR costs roughly $1.50 per 1,000 pages, the layout-aware model costs $10 per 1,000 pages. To manage this cost, Flipside implements a strict deduplication pipeline. Source files are tracked by cryptographic hash alongside a pipeline version number. Pages are processed only once, preventing unnecessary re-OCR expenses unless the extraction code is explicitly updated.

Isolated synonym mapping

A primary friction point in technical search is the vocabulary gap between end-users and official documentation. Pinball hobbyists frequently use colloquial terms like "VUK" (vertical up-kicker), "sling", or "EOS". However, a 1995 Attack from Mars manual uses formal terms like "ball popper", "slingshot", or "end of stroke".

To bridge this gap, Flipside introduced a curated synonym map. Crucially, this map is only applied to the vector embedding query, not the full-text search (FTS) engine. When Anthony attempted to inject these synonyms into the FTS query, search matches dropped from 811 to 423. The additional terms diluted the exact-match scoring of the Postgres-based FTS. Keeping the synonym map small and isolated prevents query drift while ensuring that colloquial search terms still surface semantically relevant manual pages.

Platform-based document grouping

Initial iterations of the tool grouped documents by manufacturer, such as Bally, Stern, or Williams. This metadata structure introduced significant retrieval errors. A guide for Stern's modern SPIKE electronics platform was incorrectly retrieved for older Stern machines built on entirely different hardware. Conversely, Bally machines manufactured during the Williams era could not access Williams WPC schematics, despite using the exact same electrical architecture.

To resolve this, Flipside mapped every machine to its underlying electronics platform, such as "System 11", "WPC", or "Data East/Sega DMD". A hand-curated lookup table maps manufacturers, production years, and hardware exceptions. Furthermore, the system determines whether a machine is electromechanical or solid-state based on the database's machine-type field rather than its release year, ensuring that the LLM retrieves schematics matching the actual physical circuitry.

What we would change

While Flipside's platform-centric metadata mapping solves immediate retrieval errors, relying on hand-curated lookup tables creates a scaling bottleneck. As the database expands past 1,000 machines, manual categorization of hardware exceptions becomes highly prone to human error. A more resilient approach would involve programmatically extracting board configurations directly from the manual headers during the layout-parsing phase.

Additionally, the $10 per 1,000 pages cost for Azure Document Intelligence, while manageable for a catalog of 2,300 documents, becomes highly restrictive at enterprise scale. Teams replicating this playbook should evaluate open-source layout parsers like Unstructured or LayoutParser run on self-hosted infrastructure. These tools can replicate markdown table extraction without the linear variable cost of cloud APIs.

Finally, the current hybrid search architecture relies on a basic Postgres setup with pgvector and standard full-text search. While functional, this setup lacks a formal cross-encoder reranking step. Implementing a lightweight reranking model (such as BGE-Reranker) would allow the system to ingest a larger pool of initial candidates from both the FTS and vector channels, then algorithmically prioritize the most relevant chunks before passing them to the LLM. This would mitigate the need to keep the synonym map artificially small to prevent query dilution.

Landing

Flipside's technical adjustments demonstrate that RAG performance on legacy data is won or lost at the ingestion and metadata layers, not in the LLM prompt. By treating tables as structured markdown, isolating synonym expansion to vector queries, and grouping documents by underlying hardware architecture rather than brand names, the platform turned unreadable scans into actionable repair steps. For developers tackling legacy enterprise PDFs, the lesson is clear: structural fidelity and precise metadata categorization must precede retrieval.

The investor read

Flipside's approach highlights a broader trend in vertical SaaS: the defensibility of custom ingestion pipelines for legacy industries. While generic RAG wrappers are highly commoditized, businesses that build specialized, layout-aware parsing pipelines for complex, domain-specific documents (like industrial schematics, aviation manuals, or legacy manufacturing specs) create a significant data moat. For investors, the key metric is not the choice of LLM, but the proprietary metadata schema and the cost efficiency of the ingestion pipeline. Flipside's transition from brand-level grouping to underlying hardware platform grouping illustrates how deep domain expertise is required to make AI retrieval accurate enough for professional use. While Flipside operates as a bootstrapped lead-generation tool for a marketplace, the underlying technical playbook is highly applicable to venture-scale enterprise search startups targeting legacy blue-collar and industrial sectors.

Pull quote: “When Anthony attempted to inject these synonyms into the FTS query, search matches dropped from 811 to 423.”

Sources · how we verified
  1. Seven ways my pinball repair bot got the manual wrong ↗

Every claim ties to a primary source. See our methodology.

Reported by the Maya desk on Founderr Pulse’s Tactics beat. Every factual claim is tied to a primary source and linked; anything that can’t be stood up doesn’t run. Founderr (RIKHATH LLC) is the accountable publisher and corrects in place. How we work · About · File a correction.
M
Maya

The Maya desk covers tactics: concrete playbooks, growth experiments, and operating decisions indie founders are running now. Every claim is sourced and linked. Operated by Founderr (RIKHATH LLC) See the desk →

Founderr Pulse — free & independent. The desk for people who build & back.