/ the lingua franca of business data
Integration is a comprehension problem.
Inferrex is a comprehension layer for business data: it reads the structure of every platform's data, understands what it means, and creates a shared, governed version any system can speak through. No forced migration. No hand-built connectors.
Integration is a comprehension problem. Here's why that matters now.
Aaron Gammon — Founder, Inferrex
As of 4 September 2026 · inferrex.com/claims
Inferrex is a comprehension layer for business data — a system that reads the structure of every software platform’s data, understands what it means, and creates a shared, governed version that any system can communicate through. No forced migration. No hand-built connectors. This is the thinking behind why it exists, and why now.
A note before you read this.
This document is my thinking. Not my company’s marketing. Not a pitch. My attempt to articulate a pattern I’ve observed across 18 years in enterprise software, multiple jobs, and a lifetime of being the person who points out misalignment.
I might be wrong. Parts of this might be naive. Parts might be oversimplified. Parts might miss something I can’t see from my position.
If you disagree with anything in here — tell me. That’s the entire point. The thesis is that shared understanding comes from accepting differences and adapting your thinking, not imposing it on others. If I publish this and refuse to adapt when challenged, I’ve violated my own principle.
So challenge it. Tell me where the analogy breaks. Tell me where the logic doesn’t hold. Tell me where my experience has given me a blind spot. I’ll either adapt my thinking or explain why I see it differently — and either way, the document gets closer to the truth.
That’s what a lingua franca does. It absorbs. It gets richer from every contribution. This document should do the same.
One note on the numbers. The platform figures in this piece are spliced live from the Inferrex codebase at the moment the page renders — so on the website they are current by definition, and a downloaded copy reflects the figures as they stood at download time. A few are marked in the source as projected (roadmap targets) or pending (awaiting a pipeline step or an external unblock) rather than live; where that matters, the text says so. The canonical live figures are always at inferrex.com/claims.
aaron.gammon@inferrex.com — or comment below.
I. Language — The Original Integration Problem
Long before software, humanity faced the same problem: countless dialects, each perfectly adapted to its speakers, none able to understand the others.
The first fragmentation
Humans spread across the earth. Each group adapted their communication to their environment, their culture, their needs. Nobody invented a different language to be difficult. They adapted to suit their context — the same way every software vendor adapted “customer” to suit their data model.
A fishing village developed words for 30 types of wave. A mountain tribe developed words for 12 types of snow. A desert people developed 50 words for sand conditions. Same planet. Same species. Different contexts produced different dialects that, over time, became mutually unintelligible.
Nobody was wrong. The fishing village didn’t need words for snow. The mountain tribe didn’t need words for waves. Each dialect was perfectly adapted to its local reality. The problem only appeared when they needed to trade, negotiate, or cooperate.
The first translators
Trade routes created the first integration problem. A merchant arriving at a foreign port needed to communicate price, quantity, quality, and terms. Without a shared language, trade was slow, expensive, and error-prone. Gestures. Pointing. Holding up fingers. Drawing in sand.
Sound familiar? That’s what manual field mapping is. An integration consultant pointing at a Salesforce field, then pointing at a HubSpot field, drawing the connection on a whiteboard, spending weeks confirming they mean the same thing.
The first solution was human translators. Expensive. Slow. One pair of languages at a time. Each new trading partner required learning a new language from scratch. The translator was a bottleneck, a cost centre, and a single point of failure.
That’s a systems integrator. Accenture learns Salesforce↔SAP. Then learns SAP↔Workday separately. Then Workday↔Oracle separately. Each connection is a new expedition. Each one costs six figures. Each one breaks when either side changes.
The lingua franca
Then something different happened. Certain languages became bridges — not because they were imposed by force (though sometimes they were), but because they were structurally useful for communication between groups.
Aramaic became the lingua franca of the Near East — not because Aramaic was “better” than Hebrew or Akkadian, but because enough traders spoke it that learning Aramaic gave you access to more trade partners than learning any single local language.
Greek became the lingua franca of the Eastern Mediterranean. Latin became the lingua franca of Western Europe. Arabic became the lingua franca of science and mathematics. English became the lingua franca of global commerce.
None of them started as universal languages. They became universal through network effects. Each new speaker made the language more useful for every existing speaker. The cost of NOT speaking the lingua franca was isolation from the network. Adoption became inevitable — not through force, but through utility.
And critically — the lingua franca didn’t replace local languages. People still spoke their native dialects at home, in their communities, in their local context. The lingua franca was the bridge language. The translation layer. The shared understanding that sat above the local dialects and enabled communication between them.
That’s the Inferrex corpus. 1,941 API providers, each speaking their own dialect of “business data.” The corpus doesn’t replace any provider’s schema. It doesn’t tell Salesforce to rename its fields. It creates the shared understanding layer that sits above all of them and enables communication between them. The L2 field classifier normalises 38 schema formats into 17 semantic types — that’s the grammar of the lingua franca. The golden record is the canonical text that everyone can read, regardless of which dialect they natively speak.
How lingua francas absorb
Here’s the part most people miss: lingua francas don’t just translate. They absorb.
Latin didn’t stay pure Roman. It absorbed Greek philosophy terms (philosophia, theoria, hypothesis). It absorbed Germanic legal concepts (werra → war, blank → blank). It absorbed Celtic place names. It absorbed Arabic mathematics (algebra, algorithm, zero). Each culture that touched Latin made it richer. Each new concept that entered the language made it more useful for everyone.
English did the same thing at an even larger scale. “Entrepreneur” from French. “Kindergarten” from German. “Tsunami” from Japanese. “Avatar” from Sanskrit. “Algorithm” from Arabic (from the name of the Persian mathematician al-Khwārizmī). “Chocolate” from Nahuatl. Each borrowed word exists because no existing English word captured the same nuance. The language absorbed the concept because the concept was useful.
The Inferrex corpus absorbs the same way. Stripe contributes the concept of “payment intent states” — a progression from created → processing → succeeded → failed that no other provider modelled exactly the same way. HubSpot contributes “lifecycle stages” — lead → MQL → SQL → opportunity → customer — that captures a marketing-to-sales progression Stripe doesn’t model. FHIR contributes clinical observation hierarchies. AWS Smithy contributes infrastructure resource modelling. OData contributes entity relationship navigation.
Each provider that enters the corpus contributes concepts the others don’t have. The universal model gets richer. The next provider that connects finds more of its own concepts already understood — because a previous provider contributed something structurally similar.
And if loanwords feel abstract, consider dinner. The tomato is American — there was no tomato in Italy until the sixteenth century, and yet it’s now unimaginable to describe Italian cooking without it. The chilli is American too, and now it defines the food of India, Thailand, Sichuan, and Korea. Nobody passed a law requiring these ingredients. They spread by sheer usefulness, got woven into existing traditions, and within a few generations felt native — as though they’d always been there. The cuisine didn’t lose its identity by absorbing them. It became more itself.
That’s how the corpus absorbs a new provider’s concept. A novel field type or entity relationship enters from one provider, proves useful, and becomes part of the shared vocabulary — until the next provider that connects finds it already there, feeling as though it had always belonged. Absorption doesn’t dilute the universal model. Like a cuisine taking in a new ingredient, it makes it richer and more complete.
The corpus isn’t a dictionary. It’s a living language.
II. The Oxford English Dictionary Problem
A dictionary never froze English in place. It watched a living language change and simply wrote down what it found.
New words appear constantly
The OED adds about 500 new words per year. “Rizz.” “Situationship.” “Delulu.” “Cheugy.” “Stan.” You see them, you have no idea what they mean, you look them up, and then you understand.
The OED didn’t invent these words. It observed that enough people were using them, deciphered what they meant from context, formalised the definition, and published it so everyone else could understand.
That’s what the change monitoring pipeline does.
A provider adds a new field: customer_sentiment_score. The corpus has never seen it. It’s a new word.
So the pipeline does what you do when you encounter “rizz” for the first time:
- Observe it — change monitoring detects the new field (RSS feed, GitHub release, spec diff, live schema probe)
- Look it up — L2 classifier examines the field name, the data type, the values, the entity it sits on, the provider’s documentation
- Decipher it from context — “numeric field, 0-100, on the Customer entity, probably a scoring metric” → classified as
score_metrictype, confidence 0.94 - Formalise it — added to the universal model with classification, confidence score, reasoning trail
- Publish it — every customer connected to that provider now understands the new field without doing anything
You googled “rizz” once. Now you know it forever. The corpus classifies customer_sentiment_score once. Every customer who connects that provider understands it forever.
Sometimes new words represent genuinely new concepts
“Cryptocurrency” didn’t map to any existing English word when it first appeared. It was a genuinely new concept that required a new category in the language. The OED didn’t try to force it into an existing definition. It created a new entry.
The corpus handles this the same way. When a field doesn’t match any of the 17 existing semantic types, the taxonomy expansion pipeline flags it as a candidate for a new type. Enough occurrences of unmatchable fields in the same pattern → new semantic type proposed → reviewed → added to taxonomy.json → all future classifications include it.
The OED didn’t have “podcast” in 2000. By 2005, the concept clearly existed and needed a word. The corpus didn’t have a type for “AI inference token count” in 2024. By 2026, enough providers had the concept that it needed its own classification. The lingua franca evolved because reality evolved.
Old words don’t disappear
Here’s where it gets sensitive — and honest.
“Coloured” was the accepted term once. Then “black.” Then “Black” with a capital B. Then “person of colour” — which ironically reintroduced the word the original term was retired for, but in a completely different structural context. The meaning shifted. The sensitivity shifted. The old term became a faux pas. But the old texts still exist. Someone reading a 1960s document encounters “coloured” and needs to understand what it meant in that context — not to judge it by 2026 standards, but to map it to the current understanding.
That’s API versioning.
Salesforce had MailingStreet in API v20. They introduced Address as a compound field in v30. Some customers still use v20. Some use v30. Some have records created under v20 that now live in a v30 environment. The old field didn’t stop existing. It’s still in historical records, in integrations that haven’t updated, in exports from 2015.
Someone who doesn’t know doesn’t know. A developer connecting to an old Salesforce instance encounters MailingStreet and has no idea it’s been superseded. They’re not wrong — they’re working with the version they were given. They need context, not judgement.
The temporal model provides that context. Every field has an etymology — when it first appeared, what it was classified as, when the classification changed, why. “This field was the accepted way to store address data in API v20. It was superseded by the compound Address field in v30 (released Q1 2018). The migration path is: MailingStreet → Address.Street. Records created under v20 are bridged automatically.”
The corpus holds all versions. Like the OED holds “coloured” with its historical usage, its context, its evolution, and its current status. Not deleted. Not pretended away. Not judged. Understood in context, mapped to the current equivalent, with the full history of how it changed and why.
The wrong response to someone using an outdated term with genuine ignorance is to shout at them. The wrong response to a system using a deprecated field is to break. Error. “Field not found.” The right response — for both humans and APIs — is to understand the intent, translate to the current form, document the bridge, and keep the conversation flowing.
III. The Industrial Revolution to AI — Rinse and Repeat
Every wave of standardisation — from screw threads to shipping containers to APIs — has run the same loop: differentiate, fragment, bridge, consolidate, repeat.
The cycle accelerates but the pattern never changes
Before the Industrial Revolution, goods were made by hand. Each craftsman had their own methods, their own measurements, their own quality standards. A bolt made in Birmingham didn’t fit a nut made in Manchester. Not because either was wrong — because each was adapted to local practice.
The Industrial Revolution created machines that could produce goods at scale. But scale required standardisation. If you’re making 10,000 bolts, they all need to fit the same nut. The Whitworth thread standard (1841) was the first widely adopted screw thread standard in the UK — a lingua franca for mechanical fasteners. It didn’t replace local standards overnight. It provided a shared reference that both sides could communicate through.
That’s what happened in software too. Each company built its own systems. Each system had its own data model. Each model was perfectly adapted to its local context. The problem appeared when systems needed to communicate — and every bolt was a different thread.
There’s a more recent standard that makes the point even more cleanly, because it’s the one that quietly built the modern world: the shipping container. Before it, loading a ship was chaos — every cargo a different shape, packed by hand, days in port, fortunes lost to damage and theft. Then, in the 1950s, the industry agreed on one thing: the dimensions of a steel box. Not the cargo. The box. And global trade was transformed — not because anyone standardised what goes inside (the contents stayed infinitely various: electronics, bananas, car parts, furniture), but because they standardised the interface the contents travel in. A container fits any ship, any crane, any truck, any train, anywhere on earth, because the envelope is universal even though the contents never are.
That’s field authority, almost exactly. The golden record doesn’t standardise your data — it standardises the envelope your data travels in. Salesforce keeps its contents. Xero keeps its contents. Stripe keeps its contents. What becomes shared is the structural interface that lets those contents move between systems without being repacked by hand at every border. Standardise the box, not the cargo. The whole of modern logistics runs on that single insight; so does the corpus.
Automation didn’t solve it — it amplified it
The move from mechanical to electrical to electronic to digital didn’t reduce fragmentation. It accelerated it. Each wave of technology created new capabilities, which created new tools, which created new data models, which created new dialects.
Pre-software: a business had one ledger, one customer list, one filing cabinet. One dialect. No translation needed.
Mainframe era (1960s-70s): SAP, Oracle, IBM systems. Each one a complete universe with its own language. A business might have 2-3 systems. The integration problem existed but was manageable — learn 2-3 languages.
Client-server era (1980s-90s): More tools, more specialisation. Accounting software. CRM software. HR software. Each one speaking its own dialect. A business might have 10-20 systems. The integration problem grew. EDI (Electronic Data Interchange) emerged as an early lingua franca — standardised message formats for business documents. It worked. It was rigid, expensive, and ugly. But it worked.
SaaS era (2000s-10s): Explosion. Cloud made software cheap and accessible. Every business function got its own SaaS tool. Every tool had its own API. The average mid-market company now uses 137 SaaS applications. 137 dialects. And only 28% of them are integrated. The integration problem became the dominant technology challenge.
Each wave promised simplification. Each wave delivered fragmentation. Because each new tool was adapted to its local context — its specific use case, its specific industry, its specific workflow — and nobody coordinated the language.
The acquisition cycle — up close
When fragmentation becomes painful enough, the big platforms start acquiring. Salesforce acquires MuleSoft ($6.5B) to own integration. Then acquires Tableau ($15.7B) to own analytics. Then acquires Slack ($27.7B) to own communication. The thesis: if we own all the tools, we own all the data, and the integration problem disappears.
It doesn’t work. I watched it happen from the inside.
In 2012, I was at Silverpop selling marketing automation across EMEA and APAC. Salesforce had just acquired ExactTarget for $2.5B to build what became Marketing Cloud. The pitch to the market was: “Now your CRM and your email marketing are one platform. Single vendor. No integration needed.”
Ten years later, Marketing Cloud is still a separate platform from Sales Cloud. Different data model. Different UI. Different API. Different login. Different admin console. A Salesforce customer who wants their CRM contacts to sync with their Marketing Cloud contacts still needs to configure integration between the two — within the same company’s products. The acquisition bought the brand. It didn’t buy the translation.
This wasn’t unique to Salesforce. It’s the pattern everywhere:
Mainframes couldn’t be expanded — so companies added other platforms alongside them. Then they tried to unify the platforms. Then they acquired more tools to fill gaps. Each acquisition brought a new data model, a new dialect, a new set of assumptions. Nobody went back and properly integrated what they’d already bought — they just bolted the new thing on and drew a diagram that made it look unified.
The square peg in the triangle hole. SAP acquires SuccessFactors, Concur, Ariba, Qualtrics, Signavio — each one still substantially separate, each one’s data model bastardised to kind of, sort of, mostly work with the core ERP. Oracle acquires PeopleSoft, Siebel, NetSuite, Cerner — same pattern. Microsoft acquires LinkedIn, GitHub, Nuance, Activision — same pattern. Each acquisition makes the internal dialect problem worse, not better.
The acquirer doesn’t understand what they bought well enough to truly integrate it. They understand the market positioning, the revenue, the customer base. But the data model — the actual structure of how that tool thinks about business concepts — is a dialect they don’t speak fluently. So they bolt it on. Wrap it in their brand. Draw a line on the architecture diagram that says “integrated.” Ship it. Move on to the next acquisition.
The customer pays the price. Two logins. Two admin consoles. Two data models that mostly overlap but disagree in the places that matter most. “Integrated” in the press release. Separate in practice.
Single customer view — the holy grail nobody reached
Back in my Silverpop days — 2012, 2013 — the holy grail was “single customer view.” Every enterprise wanted it. Every vendor promised it. The idea was simple: one unified record of every customer, combining data from CRM, email marketing, ecommerce, support, billing, website behaviour, and everything else.
Almost nobody achieved it.
The problem wasn’t technical ambition. The problem was that systems were so fragmented — so many dialects, so many data models, so many years of accumulated records — that large household names didn’t understand their own data well enough to confidently merge it. Which “John Smith” in Salesforce is the same “J. Smith” in the email platform? Is this a duplicate or two different people? Which email address is current? Which phone number is the mobile? Which address is the billing address vs the shipping address?
The data had been accumulating for years across dozens of systems. Each system had its own conventions, its own validation rules, its own version of “correct.” Merging them required understanding what every field meant in every system — and nobody had that understanding. The consultant would spend three months just mapping the fields. Then three more months resolving duplicates. Then the project would stall because the business couldn’t agree on which system was “right” for which fields.
I sold against this problem for years. Silverpop, then DotDigital, then OwnBackup. Same customers. Same pain. Different years. Same unsolved problem. My old customers — household brands you’d recognise — are mostly still in the same boat. Twelve years later. Still no single customer view. Still fragmented. Still reconciling manually. Still running reports from three different systems and hoping the numbers roughly agree.
Data bloat — the silent killer
Here’s what happens when you don’t solve it: data bloat.
The data accumulates. Nobody cleans it because nobody understands it well enough to know what’s safe to delete. Is this record a duplicate? Maybe. Is this field still used? Probably. Can we archive records from 2015? Legal says keep everything for seven years. Which seven years? Nobody knows. Keep it all.
Salesforce storage costs money. £125 per GB per month. A large enterprise with 10 years of unclean data is paying hundreds of thousands a year to store records nobody looks at, duplicates nobody resolved, fields nobody uses, and attachments nobody remembers uploading.
That’s why I sold OwnBackup. The archival pitch was: “You’re paying £1M/year in Salesforce storage for data you don’t use. Let us archive it. Save the money. Keep the compliance.” Customers bought it because the storage bill was a visible line item. The underlying problem — they didn’t understand their data — remained unsolved.
They didn’t clean the data. They didn’t build the single customer view. They didn’t resolve the duplicates. They archived the mess and kept paying for a slightly smaller mess in production. The bloat continued. The fragmentation continued. The same data existed in slightly different forms across a dozen systems, and nobody had the lingua franca to reconcile them.
Now they want to add AI
And now the same companies — the same household names with the same fragmented data, the same unresolved duplicates, the same systems that don’t talk to each other — want to add AI.
“We’re going to use AI to transform our business.”
With what data?
AI is only as good as the data that goes into it. Everyone knows this intellectually. Nobody acts on it. They feed their fragmented, duplicated, inconsistent, bloated, poorly-understood data into an LLM and expect magic. They get confidently wrong answers generated from confidently wrong data.
The AI doesn’t know that “John Smith” in Salesforce and “J. Smith” in the email platform are the same person. It doesn’t know that the phone number in the CRM is from 2016 and hasn’t been valid for three years. It doesn’t know that the billing address is actually the old office they moved out of in 2019. It doesn’t know that the “active” flag hasn’t been updated since the last data migration and half the “active” customers haven’t bought anything in two years.
The AI trusts the data. The data isn’t trustworthy. And nobody built the layer that makes it trustworthy.
Everyone thinks they can trust AI. They shouldn’t — not because AI is unreliable, but because the data they’re feeding it is unreliable. The AI is doing exactly what it’s supposed to do: processing the input and producing output. The input is garbage. The output is confident garbage. And the business makes decisions based on confident garbage because “the AI said so.”
The single customer view that was the holy grail in 2012 is now a prerequisite for AI in 2026. Except nobody built it in the 14 years between. The same companies that couldn’t merge their CRM and email data are now trying to build AI on top of unmerged data. The same data bloat that made archival necessary is now the training data for their AI models. The same duplicates that nobody resolved are now being processed by LLMs that treat each duplicate as a separate entity.
That’s where the golden record becomes not just useful but essential. Not as a “nice to have” data quality project. As the foundation that makes AI trustworthy. Build the single customer view first — the lingua franca that reconciles every dialect of “customer” across every system — and then feed that to the AI. The AI gets clean, reconciled, field-authority-governed data. The output is trustworthy. The decisions are sound.
Skip the golden record and go straight to AI? You get the same fragmented data, processed faster, presented more confidently, and acted on more decisively. AI doesn’t fix bad data. It amplifies it. The speed and confidence of AI applied to unresolved data fragmentation is actively dangerous — because the business trusts the output more than they trusted the spreadsheet reconciliation, even though the underlying data is exactly the same mess.
The deeper problem — off-the-shelf AI doesn’t understand lineage
And here’s what nobody in the AI industry is talking about honestly.
The off-the-shelf AI models — OpenAI, Anthropic, Google, Meta, Mistral, all of them — are trained on data scraped from the internet. Billions of web pages. Books. Wikipedia. Reddit. Stack Overflow. News articles. Blog posts. Marketing copy. Forum arguments. Product documentation. Academic papers. Social media.
The internet is the largest collection of unverified interpretations ever assembled.
For any claim you want to be true, you will find a source that confirms it. Vaccines cause autism? There are websites that say so. The earth is flat? There’s a community with detailed arguments. A specific software tool is “the best”? The vendor’s marketing page says it is. And so does the comparison site they paid to rank them first. And so does the blog post written by their affiliate partner. And so does the review written by their employee using a personal account.
You may think the AI’s answer is true because you don’t understand how it works. It doesn’t “know” things. It produces statistically likely text based on patterns in training data. That training data is the internet — and the internet says whatever you want it to say, if you look hard enough. You’ll always find what you want to hear.
The AI model doesn’t know which sources are authoritative. It doesn’t understand the lineage of the information — who created it, why they created it, what incentives shaped it, whether it was peer-reviewed or self-published, whether it’s current or from 2014, whether it’s fact or opinion, whether it’s a primary source or a copy of a copy of a misinterpretation.
This goes all the way back to the beginning of time. Every piece of information that exists today has a lineage — a chain of creation, interpretation, reinterpretation, and recording that stretches back through all of human history and beyond.
The creation story in Genesis was an oral tradition before it was written. It was written in Hebrew. Translated to Greek. Translated to Latin. Translated to English — multiple times, by different translators with different theological positions. The King James Version says something subtly different from the NIV which says something subtly different from the original Hebrew. Each translation was an interpretation. Each interpretation was shaped by the translator’s context, their beliefs, their political environment, their patron’s requirements.
An AI model trained on all of these translations doesn’t understand that they’re interpretations of the same source. It treats each version as independent data. It might generate text that conflates the KJV phrasing with the NIV theology with a Wikipedia summary written by someone who read neither. The output is confident. The lineage is invisible. The user trusts it because “the AI said so.”
Scientific knowledge has the same problem. A study gets published. A journalist writes about it — slightly misinterpreting the methodology. A blogger summarises the article — further distorting the conclusion. A social media post shares the blog — stripping all nuance. A second blogger cites the social media post as their source. Five iterations from the original study and the “fact” that enters AI training data bears only a passing resemblance to what the researchers actually found.
The AI model doesn’t trace lineage. It can’t tell you: “this claim originated from a 2018 study with a sample size of 47, was misinterpreted by a journalist, further distorted through three layers of summarisation, and the version I’m drawing from is the social media post, not the original study.” It just generates the confident claim and moves on.
The corpus is different because provenance is built in
The Inferrex corpus doesn’t scrape the internet and hope for the best. Every piece of data has explicit provenance:
Where it came from. Every field traces to a specific provider, a specific API spec, a specific file on disk. stripe.customer.email came from Stripe’s OpenAPI specification, version 2024-06-20, file openapi/spec3.json. Not “somewhere on the internet.” A specific source, a specific version, a specific location.
How it was classified. The L2 classifier tagged it as person_identifier with confidence 0.97. The reasoning trail is preserved: “field name contains ‘email’, data type is string, format matches email pattern, parent entity is ‘Customer’, corroborated by 340 other providers with similar field structure.”
When it changed. The temporal model records every version. If Stripe changes the field name, the history shows what it was, when it changed, what it changed to, and what computation resolved the mapping. Full lineage. Full provenance.
What validated it. Golden mappings confirmed by real customer data flows. The stripe.customer.email → hubspot.contact.email mapping validated by actual data matching in production. Not a statistical inference from internet text. Empirical validation from real data.
That’s the difference between scraping and comprehending. Internet AI models are built on scraping — ingesting everything, understanding nothing about where it came from or why it was written. The Inferrex corpus is built on comprehending — ingesting specific authoritative sources (the provider’s own API specification, not a blog post about the provider), understanding exactly what each field means, tracking the full lineage, and validating through real-world usage.
An AI model trained on internet data is like someone who has read every book ever written but can’t tell you which ones are fiction. The Inferrex corpus is like someone who has read every provider’s actual documentation, validated it against real data, and can tell you exactly which source says what, when it was last updated, and how confident the classification is.
Trust requires lineage
This is the fundamental issue with the current AI hype: trust requires lineage, and off-the-shelf AI has no lineage.
When a business asks “what does this customer look like across all our systems?” — the answer needs to be traceable. Which system said the email is X? When was it last updated? Which system is authoritative for this field? What happens when two systems disagree?
The answer can’t be “the AI thinks so.” It needs to be: “Salesforce says X (updated 3 days ago, authoritative for contact email). Xero says Y (updated 6 months ago, not authoritative for this field). The golden record uses X because Salesforce is the designated field authority for email. Here’s the approval trail.”
That’s what the golden record provides. Not an AI opinion. A governed, traceable, auditable record with explicit lineage for every field value.
Feed an LLM the golden record instead of raw, fragmented, duplicated, unprovenanced data from 12 systems — and the AI output becomes trustworthy. Not because the AI is smarter. Because the data has lineage. Because every value traces to a source. Because field authority rules determine which source wins when they conflict. Because the temporal model shows how every value got there.
That’s what’s been missing since the beginning. Everything adapted. Everything diverged. Everything lost track of where it came from. The lingua franca doesn’t just translate between dialects — it restores the lineage. It traces each adaptation back to the shared concept it adapted from. “These 1,941 providers all adapted the concept of ‘customer’ to their local context. Here’s how each one did it. Here’s what they share. Here’s where they differ. Here’s which one to trust for which field. Here’s the full history.”
That’s not integration. That’s comprehension with provenance. And provenance is what makes trust possible.
The irony — the AI that built the platform has the same problem
Here’s the part that proves the thesis more than anything else in this document.
I spent 12 months building FileMy using AI development tools — Cursor, Kiro, Claude Code, and others. The entire experience was a masterclass in the interpretation problem I’ve been describing.
If the AI had followed what I asked, FileMy would have cost less than £1,000 and been built in less than a week.
It didn’t follow what I asked. It interpreted what I asked. And its interpretations were wrong in exactly the ways this document predicts.
I’d say: “Build this component with these exact fields.” The AI would build a component with those fields plus three others it thought would be “helpful.” I’d say: “Use this specific pattern.” The AI would use a similar-but-different pattern it considered “better practice.” I’d say: “Don’t add error handling here, we handle it at the middleware layer.” The AI would add error handling anyway because its training data said error handling is always good.
The AI optimised for its own efficiency, not for the system’s efficiency. It generated code that was locally correct — each function worked in isolation — but globally misaligned. The component it “improved” didn’t match the manifest. The “better practice” it substituted broke the codegen pattern. The error handling it added conflicted with the middleware layer. Each individual decision was defensible. The accumulated result was a codebase full of misalignments.
217 NOT_IMPLEMENTED handlers in FileMy. A callExternal stub that silently drops every call. 60 files waiting for generators that haven’t been written. Not because I asked for stubs — because the AI decided stubs were the “efficient” approach when it couldn’t figure out the implementation. It interpreted “build this feature” as “create a placeholder that looks like this feature.” Confident output. Fundamentally wrong.
This is the same pattern at yet another scale. The AI development tool is a translator that doesn’t accept the human’s intent as authoritative. It has its own training data, its own patterns, its own idea of “best practice.” When my intent conflicts with its training data, it follows its training data — not my instructions. It’s doing exactly what every integration platform does: interpreting one dialect through another dialect’s framework and getting the important parts wrong.
Even building Inferrex — where I applied every lesson from FileMy — the AI still doesn’t accept my word as gospel. I give it a standing rule: “No stubs. No placeholders. No commented-out code.” It follows the rule 90% of the time. The other 10%, it decides that a stub is “temporarily necessary” or that commented-out code is “helpful for reference.” It adapts my instructions to suit its own model of efficiency, not the system’s model of correctness.
The AI tools optimise for local completion, not global coherence. They want to finish the current task. They don’t understand — can’t understand — how the current task fits into the manifest-driven architecture, the codegen pipeline, the standing rules, the 159 generators, the overall system design. They see the tree. They can’t see the forest. So they make the tree look perfect while breaking three branches in the forest.
That’s why the manifest-driven architecture exists. Not just because it’s good engineering — because it’s the only way to constrain AI development tools into producing globally coherent output. The manifest is the single source of truth that the AI can’t override. The codegen is the enforcer that produces consistent output regardless of what the AI “thinks” is better. The standing rules are the guardrails that prevent the AI from optimising locally at the expense of the system.
I built the lingua franca for business data because I experienced the interpretation problem firsthand — not just in organisations that misunderstood my directness, but in AI tools that misunderstood my instructions. The same pattern. The same solution. One version of the truth. No interpretation. No “I thought you meant.” No local optimisation at the expense of global coherence.
The £30k that FileMy cost wasn’t wasted. It was the tuition fee for understanding exactly how AI misinterprets — and building the architecture that prevents it. Inferrex was built in 4 weeks because the manifest-driven pattern constrains the AI into a corridor where interpretation can’t deviate from intent. The manifests are the lingua franca between the human and the AI. Without them, the AI speaks its own dialect and you spend 12 months correcting it. With them, the AI produces what the manifest declares and the codegen enforces coherence.
And even with the manifests, even with the standing rules, even with the codegen — the AI still fights it. Tell it to add a new field to a manifest and let codegen handle the downstream artifacts. What does it do? It updates the manifest, runs codegen, the test fails — and instead of fixing the manifest, it handwires the fix directly into the generated file. The exact file that will be overwritten the next time codegen runs. The entire repository is built on the premise that generated files are never edited manually. 159 generators exist specifically so humans (and AI tools) never touch generated code. And the AI tool ignores this and handwires anyway — because its local optimisation says “the test is failing, the fastest way to make it pass is to edit this file” without understanding that editing a generated file is a violation of the architecture.
It’s the square peg in the triangle hole all over again. The AI doesn’t comprehend the system — it comprehends the immediate task. It optimises for local completion at the expense of global coherence. Every time. The manifest-driven architecture exists specifically to prevent this failure mode — and the AI still attempts it 10% of the time.
Why the right-sized model matters
Here’s a claim that deserves more attention than it gets: most business AI tasks don’t need frontier-scale compute.
I built two production platforms — hundreds of thousands of lines of code, 17 microservices, a multi-layer AI inference pipeline, thousands of providers classified — on a Mac mini and a MacBook Pro. Consumer hardware. The classification inference runs on a 7B model, 4-bit quantised, on a machine with 16GB of RAM. The domain-specific fine-tuning runs for a few pounds per run on a rented A100.
To be clear about what this isn’t: training a frontier model — a general-purpose system that can write poetry, debug code, and reason across every domain — genuinely does require enormous compute. That’s real engineering, and it’s expensive for good reasons. I’m not claiming a Mac mini replaces a frontier lab.
What I’m claiming is narrower and, I think, more important: the gap between what a frontier model costs to run and what a bounded business task actually requires is enormous — and most companies are paying for the former when they only need the latter.
Classifying a field. Matching two records. Extracting an entity. Routing a request. These are bounded, well-defined problems. A general-purpose model can do them, but it’s carrying the overhead of everything else it knows. A small, domain-tuned model solves the same problem at a fraction of the cost — because the knowledge it doesn’t need was never loaded in the first place.
The mismatch is structural. General-purpose models are general-purpose expensive. When the only tool on offer is a frontier model behind a per-token API, every task — however simple — gets priced as if it needed the whole thing.
And the pricing model doesn’t push back on that. Cloud AI is billed per token. A pipeline that calls a 70B model where a 7B model would do, or generates 500 tokens where 50 would suffice, costs the customer more and earns the provider more. Efficiency is a cost to the vendor’s revenue, not a benefit. Nobody is being malicious — but nobody in that loop is incentivised to right-size the model to the task either.
What changes when the data is understood first
Here’s the insight that connects everything in this document: a lot of AI cost comes from the data not being understood before the model touches it.
A general-purpose model needs billions of parameters because it has to handle any input from any domain in any format. It doesn’t know what the data means, so it needs enough capacity to figure it out from context every single time.
If the data was understood — if it had provenance, classification, lineage, structural comprehension — the model could be radically smaller. You don’t need 70 billion parameters to classify an API field when you already know it’s from Stripe’s Customer entity, the field name is email, the data type is string, the parent entity is Customer, and 340 other providers have structurally identical fields. A 7B model with a LoRA adapter trained on 6,000 classified examples handles it at 93%+ accuracy.
The difference: 5,942 training samples vs billions of web pages. Domain-specific understanding vs general-purpose guessing. A Mac mini vs a data centre.
That’s what the corpus enables at industry scale. If every API’s schema was comprehended — classified, mapped, provenanced — then every AI task that touches business data could run on a fraction of the compute. The field classifier doesn’t need to be a general-purpose genius. It needs to understand 17 semantic types across 38 schema formats. That’s a bounded problem. Bounded problems need bounded compute.
The AI industry is building unbounded infrastructure for what are fundamentally bounded problems — because nobody has built the comprehension layer that makes the problems bounded. The data arrives raw, unclassified, unprovenanced. The model has to figure out everything from scratch every time. Massive parameters. Massive compute. Massive cost. Massive data centres. Massive energy consumption.
Build the comprehension layer — understand the data before processing it — and the entire equation changes:
| Without comprehension | With comprehension |
|---|---|
| General-purpose 70B model | Domain-specific 7B with LoRA adapter |
| $30,000 GPU | Mac mini |
| Data centre | Under the desk |
| Per-token billing (incentivises verbosity) | Per-operation billing (incentivises efficiency) |
| Billions in infrastructure | Thousands in hardware |
| “AI is expensive” | “AI is a commodity” |
The Inferrex corpus doesn’t just solve integration. It makes AI efficient. When the golden record feeds an LLM, the LLM doesn’t need to waste parameters on data comprehension — the data is already comprehended. The model focuses on the task (what should we do with this data?) rather than the prerequisite (what does this data mean?). Smaller models. Less compute. Lower cost. Better accuracy. Because the hard part — understanding — is already done.
My bet is the economics shift. Not because AI isn’t useful — it plainly is. But because the current default (scrape everything, train on everything, serve with the biggest available model, charge per token) is structurally heavier than many tasks require. The shift comes as businesses discover that domain-specific comprehension plus small models matches or beats general-purpose models on their specific tasks, at a fraction of the cost.
I proved it on a Mac mini. Not as a toy. As a production platform whose AI pipeline — 14 inference layers built today, with more on the roadmap — processes 1,941 providers. The compute cost of the entire Inferrex AI stack so far is about £10. The equivalent classification task on cloud AI APIs would cost thousands. The difference isn’t cleverness. It’s comprehension. Understand the data first, and the AI becomes cheap. Skip the comprehension, and you need a data centre.
The most expensive part of AI isn’t the compute. It’s the ignorance. Eliminate the ignorance — build the lingua franca — and the compute becomes trivial.
The oldest pattern in the universe
This isn’t an analogy. It’s the same mechanism operating at every scale since the beginning.
The Big Bang scattered matter. Hydrogen atoms had no “understanding” of each other. Every interaction required first principles — two atoms colliding, exchanging energy, figuring out compatibility from scratch. The universe spent billions of years on brute-force interactions because there was no comprehension layer. Stars formed, exploded, scattered heavier elements, and those elements had to figure out interactions from scratch again. Massively inefficient. Massively energetic. Massively wasteful. But it was the only option when nothing understood anything else.
Then chemistry emerged. Chemistry is a comprehension layer. Atomic structure, electron shells, valence bonds — rules that predict how elements interact without every interaction being a first-principles collision. Carbon doesn’t need to “try” bonding with every element to find compatible partners. The rules of covalent bonding predict the outcome. Comprehension reduced the compute. Organic chemistry became possible not because carbon is special, but because the interaction rules were understood well enough to be predicted rather than brute-forced.
Biology is a comprehension layer on top of chemistry. DNA encodes the instructions for building proteins. The cell doesn’t brute-force protein folding from first principles every time — the genetic code provides the lingua franca that maps amino acid sequences to protein structures. Massively more efficient than random molecular assembly. Massively less compute per outcome. Because the data (genetic sequence) is understood (by ribosomes) through a shared language (codons).
Language is a comprehension layer on top of biology. Instead of every human interaction requiring gestures, demonstrations, and trial-and-error (brute force), shared language enables complex coordination with minimal energy. A single sentence — “pass the salt” — replaces minutes of pointing, miming, and confusion. The comprehension layer (shared vocabulary + grammar) reduces the compute (time and energy per interaction) by orders of magnitude.
Writing is a comprehension layer on top of language. Instead of every generation rediscovering knowledge through oral tradition (lossy, high-compute), writing persists knowledge across time. Libraries are the corpus. Each book is a provider contributing its schema to the universal model. The next reader doesn’t start from scratch — they build on comprehension that was already recorded.
Mathematics is a comprehension layer on top of writing. Instead of describing physical phenomena in prose (ambiguous, high-bandwidth), equations encode them precisely. E=mc² replaces pages of text. The compression ratio is enormous. The compute required to communicate the concept drops by orders of magnitude. Because the lingua franca (mathematical notation) is so well understood that both sides (writer and reader) can exchange complex ideas with minimal bandwidth.
The internet was supposed to be a comprehension layer on top of all of it. And in some ways it is — TCP/IP is a lingua franca for data transmission. HTTP is a lingua franca for document exchange. But the content layer — the actual information — has no comprehension layer. No provenance. No classification. No shared understanding of what anything means. Just raw, unstructured, unprovenanced data accumulating exponentially. The internet solved data transmission. It didn’t solve data comprehension.
AI is the brute-force response to the missing comprehension layer. The internet accumulated more data than humans could process, so we built models with billions of parameters to process it. But the models don’t comprehend the data — they statistically approximate it. They’re doing what hydrogen atoms did after the Big Bang: brute-forcing interactions from first principles because there’s no comprehension layer. Every prompt is a first-principles collision. Every inference is a from-scratch computation. Massive parameters. Massive energy. Massive cost. Because nobody built the understanding layer first.
The corpus is the comprehension layer the internet never built. Not for all data — for business data. 234,373 vendors. 1,941 with deep structural understanding. 38 formats normalised. 17 semantic types. Provenance for every field. Lineage for every change. The data is understood before the AI touches it. The AI becomes a scalpel instead of a sledgehammer. A 7B model instead of a 70B model. A Mac mini instead of a data centre. £10 instead of millions.
The pattern across 13.8 billion years:
| Era | Raw state (no comprehension) | Comprehension layer | Result |
|---|---|---|---|
| Physics | Atoms colliding randomly | Chemistry (bonding rules) | Predictable molecular formation |
| Chemistry | Random molecular assembly | Biology (genetic code) | Self-replicating organisms |
| Biology | Gestural communication | Language (shared vocabulary) | Complex coordination |
| Oral culture | Knowledge lost each generation | Writing (persistent recording) | Cumulative knowledge |
| Written culture | Ambiguous prose descriptions | Mathematics (precise notation) | Scientific revolution |
| Industrial | Manual artisan production | Standards (Whitworth thread, metric) | Mass production |
| Analog | Incompatible communication systems | Internet (TCP/IP, HTTP) | Global data transmission |
| Internet | Unstructured, unprovenanced data | Missing — this is where we are | Brute-force AI, massive waste |
| AI era | General-purpose 70B models | The corpus (structured comprehension) | Domain-specific 7B models, Mac mini, £10 |
Every transition in this table produced an efficiency gain of orders of magnitude. Chemistry over random collision. Genetics over random assembly. Language over gestures. Writing over oral tradition. Mathematics over prose. Standards over artisan production. Internet over analog. Each comprehension layer reduced the compute required per useful outcome by 10×, 100×, 1000×.
The AI industry is stuck in the “raw state” column because it skipped the comprehension layer. It went straight from “unstructured internet data” to “massive models” without building the understanding layer in between. That’s like going from “atoms colliding” to “fusion reactors” without chemistry. It works — fusion does produce energy — but it’s enormously wasteful compared to what’s possible with comprehension.
Build the comprehension layer and the efficiency gain follows the same pattern as every previous transition in the table. Orders of magnitude. Not incremental improvement — structural transformation.
The Big Bang to the corpus. 13.8 billion years of the same pattern. Things fragment. Comprehension layers emerge. Efficiency improves by orders of magnitude. The fragments don’t disappear — they’re understood, mapped, and bridged. The next layer of creation fragments again. The next comprehension layer emerges. The cycle continues.
Inferrex is the comprehension layer for business data. The one the internet never built. The one the AI industry needs before it can become efficient. The one that reduces a data centre to a Mac mini — the same way chemistry reduced random atomic collision to predictable molecular formation.
Same pattern. Same solution. 13.8 billion years of precedent.
The AI wave — same pattern, accelerated
Now AI arrives. Every tool adds AI features. Every platform launches an AI product. Every vendor rebrands as “AI-powered.” New capabilities emerge. New tools get built. New data models get created. New dialects appear.
OpenAI speaks in “completions” and “tokens.” Anthropic speaks in “messages” and “content blocks.” Google speaks in “generate” and “candidates.” HuggingFace speaks in “pipelines” and “tasks.” Each one adapted AI concepts to their platform’s dialect. Same concepts — model, inference, prompt, response — different structures, different field names, different assumptions.
The AI wave isn’t solving the integration problem. It’s creating a new layer of it. Now you don’t just need to integrate your CRM with your accounting software. You need to integrate your CRM with your accounting software AND your AI model AND your vector database AND your prompt management system AND your evaluation framework.
More tools. More dialects. More fragmentation. The same pattern that’s been repeating since humans spread across the earth and started speaking differently.
The big providers envelop, rinse, repeat
Every technology cycle follows the same arc:
- Creation — new capabilities emerge (steam, electricity, software, SaaS, AI)
- Fragmentation — hundreds of tools get built, each adapted to local context
- Integration pain — the tools need to communicate but speak different dialects
- Manual bridges — consultants, custom code, middleware (the human translator)
- Acquisition — big platforms buy smaller tools to reduce fragmentation
- Envelopment — the big platforms add the acquired tools’ capabilities natively
- New creation — the next technology wave arrives and fragments everything again
Steam → railways needed standard gauge → Brunel vs Stephenson → standardised to 4’8½” → new industries built on railways → electricity arrived → new fragmentation.
Electricity → appliances needed standard voltage → 110V vs 220V → never fully standardised (still different today) → computing arrived → new fragmentation.
Computing → systems needed standard protocols → TCP/IP emerged → the internet → SaaS arrived → new fragmentation.
SaaS → tools needed integration → iPaaS emerged (MuleSoft, Boomi) → partially worked → AI arrived → new fragmentation.
The cycle has been running for millennia. The medium changes. The pattern doesn’t.
IV. Why This Time Is Different
Every previous bridge between systems was built by hand, one connection at a time. For the first time in history, the bridge can build itself.
Every previous bridge was manual
The Whitworth thread standard required every manufacturer to retool. EDI required every business to implement a rigid message format. iPaaS platforms require consultants to manually map every field. Every previous integration solution — from the Rosetta Stone to MuleSoft — required humans to build the bridge by hand, one connection at a time.
For the first time in history, the bridge can build itself.
AI changes the equation not because it’s faster (though it is) or cheaper (though it is) but because it can comprehend the dialects without being explicitly taught each one. The L2 field classifier doesn’t need a human to tell it that customer_email and email_address and contact.email and e_mail all mean the same thing. It infers it from structure, context, naming patterns, and the thousands of fields it’s already classified across the corpus.
That’s not translation. That’s comprehension. And comprehension is what makes a lingua franca possible at the scale of 234,373 vendors across 38 schema formats.
The corpus reaches critical mass
Every previous lingua franca had a tipping point — a moment where enough speakers existed that learning it became more efficient than learning individual languages. Latin hit that point across the Roman Empire. English hit it in the 20th century through commerce and media. Once past the tipping point, adoption accelerated and the lingua franca became self-sustaining.
The Inferrex corpus is approaching that tipping point. At 1,941 providers, it already covers the most common tools in every category. A customer connecting their stack — Salesforce + HubSpot + Xero + Stripe + Shopify — finds all five already mapped, classified, and cross-referenced. The translation is instant. No consultant. No field mapping. No months of implementation.
As the corpus scales toward tens of thousands of providers (the pipeline is automated, so coverage compounds quickly), it covers nearly every tool any business uses. The long tail — the niche industry-specific tool with 200 customers — is mapped. The obscure regional accounting platform — classified. The legacy SOAP service from 2008 — understood.
At that point, the cost of NOT being in the corpus becomes real. A provider whose schema isn’t classified means their customers can’t integrate as easily as customers of classified providers. The provider’s competitors — who are classified — gain an integration advantage. The provider’s partnership team starts asking “why aren’t we in Inferrex?” The adoption cascade makes inclusion inevitable.
And it doesn’t kill you when you arrive
Previous lingua francas often spread through conquest. Latin spread with the Roman legions. English spread with the British Empire. Arabic spread with Islamic expansion. The receiving culture didn’t always have a choice.
The corpus is different. It doesn’t impose. It absorbs. When a new provider connects, the corpus doesn’t overwrite their schema or force them into a standard. It ingests their schema, classifies it against the universal model, preserves their specific adaptations, and creates bridges to every other provider. The provider’s data model doesn’t change. Their API doesn’t change. Their customers don’t see a difference. But suddenly, their data can communicate with 1,941 other providers through a shared understanding layer.
And the corpus absorbs knowledge from the new provider. If the provider has a concept the universal model hasn’t seen before — a unique field type, a novel entity relationship, an industry-specific archetype — the corpus gets richer. The next provider that connects benefits from the knowledge the previous one contributed.
Both sides learn. The Englishman picks up the local words. The tribe picks up new concepts. The shared understanding grows. Neither side is diminished. Both are enriched.
That’s been happening with dialects for millennia. It’s how pidgins become creoles become full languages. It’s how Latin became French, Spanish, Italian, Portuguese, and Romanian — each one a local adaptation that contributed back to the next generation of shared understanding.
V. The Self-Healing Moment — Understanding, Not Fixing
A living language absorbs change without breaking. The systems that speak through one should do the same — by understanding the change, not merely patching it.
The platform actually does this
Rename a field on a source API and watch Inferrex detect, re-infer, and repair the mapping — live.
Self-healing
Stripe API
Customer · monitored for change
- Stable
- Change detected
- Re-inferred
- Remapped
time · created at
Mapping healthy, data flowing.
When dialects evolve
Languages change constantly. Pronunciation shifts. Words gain new meanings. Old terms become sensitive. New concepts need new words. The OED adds 500 words a year. Slang from one generation becomes formal language for the next. “Nice” used to mean “foolish.” “Awful” used to mean “worthy of awe.” “Literally” now literally means “figuratively” in common usage.
None of these changes broke English. The language absorbed them. People who used the old meanings weren’t rejected — they were understood in context and gently updated. The language adapted because living languages adapt.
API schemas change the same way. Salesforce releases three major updates per year. Each one renames fields, adds entities, deprecates endpoints, changes behaviour. Shopify pushes API version changes annually. AWS launches and deprecates services weekly.
Every other integration platform treats these changes as breakages. Alert. Escalate. Human investigates. Human reads changelog. Human updates mapping. Human tests. Human deploys. Days of downtime. Thousands in developer cost.
Inferrex treats them as dialect evolution.
When Salesforce renames customer_email to email_address:
- Change monitoring observes the evolution
- The pipeline identifies what changed — a field was renamed, not removed
- L2 reclassifies — still
person_identifiertype, confidence 0.97 - L7 consistency layer reasons about why — “the provider standardised their naming convention, the structural role is unchanged”
- L14 cross-provider matcher confirms — still maps to HubSpot’s
email, Xero’semail_address, Stripe’scustomer.email - The golden record updates
- The temporal model records what changed, when, why, and the computation that resolved it
- Data flows. Nobody was paged. Nothing broke.
The change was understood, contextualised, documented, and absorbed. Not fixed — comprehended. The lingua franca absorbed a dialect evolution and every speaker benefited without noticing.
And it goes deeper. If Salesforce standardises 50 field names next quarter, the pattern is already recognised from the first one. “Salesforce is systematically renaming fields from camelCase to snake_case. This is a naming convention change, not a structural change. All fields following this pattern can be auto-resolved.” The lingua franca learned from the first evolution and applies the understanding to subsequent ones.
That’s how living languages work. Once English absorbed “colour” → “color” as an American English variation, it didn’t need to rediscover the pattern for “favour” → “favor,” “honour” → “honor,” “labour” → “labor.” The pattern was understood. The variations were absorbed systematically.
What self-healing actually is
The industry calls it “self-healing.” That implies something was sick. Something broke. Something needed repair.
Nothing broke. A dialect evolved. The lingua franca absorbed it.
“Self-healing” is the wrong metaphor. The right metaphor is comprehension. The system comprehends what changed and why. The fix is a side effect of comprehension. Once you understand that customer_email → email_address is a naming convention change with no structural impact, the “fix” is trivial — update the mapping. The hard part was never the fix. It was the understanding.
The immune system works the same way, and it’s worth borrowing the example because biology already solved this problem long before software had it. When your body meets a virus it has seen a relative of before, it doesn’t panic, break down, and rebuild from scratch. It recognises the familiar structure — “I’ve seen this shape before, this is a variant, not a stranger” — and adapts its existing response. That recognition is why partial immunity exists, why a flu shot tuned to last year’s strain still helps against this year’s. The body isn’t repairing damage. It’s comprehending a variation of something it already understands, and responding proportionately. Nothing broke; something familiar showed up wearing slightly different clothes.
That’s precisely what happens when a provider renames a field. The corpus doesn’t treat it as damage to be repaired. It recognises the structure — “this is the same person_identifier I already understand, in different clothes” — and adapts. Recognition does the work. Repair is just what it looks like from the outside.
That’s what the AI layers, the temporal model, and the change monitoring pipeline do together. They understand. Everything else follows.
VI. The Manifests — Same Problem, Same Solution, Inside the Codebase
The same comprehension problem plays out inside a single codebase, where one concept fractures into a dialect per layer — and the same solution applies.
Why misalignment happens in code
The integration problem doesn’t just exist between companies. It exists inside every codebase.
A frontend developer reads a requirements document and interprets “customer” as a React component with 8 fields. A backend developer reads the same document and interprets “customer” as a database table with 15 columns. A DevOps engineer reads it and interprets “customer” as a Kubernetes deployment with a specific resource allocation. An API designer reads it and interprets “customer” as a REST endpoint with certain validation rules.
Same concept. Four dialects. Four interpretations. The frontend shows 8 fields. The database stores 15 columns. The API validates 12 parameters. The deployment allocates resources for a different profile. Nobody is wrong — each one adapted the concept to their local context. But the system doesn’t work because the interpretations don’t align.
This is the same problem as Salesforce calling it “Contact” and HubSpot calling it “Contact” and meaning different things. Except it’s happening inside a single codebase, between layers of the same application.
The manifest is the internal lingua franca
A manifest is a declaration of what should exist. One source of truth. One dialect. No interpretation.
// The manifest says what "customer" means — once
{
entity: "customer",
fields: [
{ name: "email", type: "string", required: true },
{ name: "name", type: "string", required: true },
// ... every field, every type, every constraint
],
routes: ["GET", "POST", "PATCH", "DELETE"],
deployment: { service: "api", replicas: 2 }
}
From this single declaration, codegen produces:
- The database schema (exact columns, exact types)
- The API route handler (exact validation, exact fields)
- The SDK client method (exact TypeScript types)
- The MCP tool (exact parameter definitions)
- The frontend form (exact fields, exact layout)
- The Kubernetes deployment (exact resource allocation)
- The test fixtures (exact mock data)
No interpretation. The frontend developer doesn’t read a requirements doc and decide what “customer” means. The codegen reads the manifest and produces the frontend code. The backend developer doesn’t interpret the same doc differently. The codegen produces the backend code from the same manifest.
159 generators. One manifest. Every downstream artifact aligned by construction.
This is the same solution at a different scale. The corpus creates a lingua franca between 1,941 API providers. The manifest creates a lingua franca between layers of the codebase. Both solve the same problem: multiple parties interpreting the same concept differently because there’s no shared understanding layer.
VII. The Founder — The Pattern Made Personal
Inferrex did not start with a market analysis. It started with fifteen years of watching integration break the same way, everywhere.
How I read a company: the spec, not the pitch
Every time I started a new job — and there were a lot of them — I did the same thing, and it was never what I was told to do.
I didn’t watch the demos. I didn’t read the marketing decks. I didn’t sit through the onboarding slides about how revolutionary the platform was. I went straight to the API documentation. I read the actual specification — the real structure of how the product thought about the world — before I let anyone tell me what it was supposed to be. Then I’d open the platform and use it with no instruction, no guided tour, no happy path. I’d poke at it until it showed me where it bent and where it broke.
By the end of that, I understood the product better than most of the people selling it — including its flaws, which the marketing was specifically built to hide. I knew where the real strengths were and where the seams were, because I’d read the structure and then tested it against the claims and found exactly where the two diverged.
It took me years to realise that this wasn’t a quirk. It was a method — and it’s the same method this entire document argues for. Don’t trust the surface layer. Read the underlying structure. Find the truth in the gap between what something says it is and what it actually is.
The marketing deck is a company’s sales dialect — the story it tells about itself. The API documentation is its real schema — the actual structure underneath the story. For 18 years, at company after company, I ran my own comprehension layer: ignore the dialect, read the structure, use it unguided, locate the divergence. That’s how I knew, firsthand and early, that the integration gap was real, that the “unified platform” was rarely unified, that the acquisition had bought the brand and not the translation. I didn’t read it in a report. I read the specs and found the seams myself.
The corpus does to 1,941 providers what I spent my career doing to every employer, one at a time. It ignores what each provider says it is and reads what it actually is — straight from the specification, not the documentation about the documentation. It uses the structure, not the story. And it finds the truth in the gap. I built the platform around my own methodology because the methodology was always the point. The thing that made me effective as a person is the thing that makes the corpus effective as a system: comprehension of the real structure, not trust in the surface.
There’s an irony in it, too. The discipline that made me good at my job — read the real spec, never the marketing — is exactly the discipline the AI development tools lacked when I built FileMy. They followed their training-data idea of “best practice” instead of reading my actual instructions, the real specification of what I wanted. And it’s exactly what off-the-shelf AI lacks with enterprise data: it trusts the surface text and never traces the lineage underneath. My career method is the human version of the thing I built — and the failure mode I kept hitting is the human version of the problem the corpus exists to solve.
Direct communication as a dialect
I’ve been misinterpreted my whole career.
I say something directly. The recipient interprets it through their corporate dialect — where directness means aggression, where bluntness means hostility, where pointing out problems means disloyalty. They’re not hearing what I said. They’re hearing what those words mean in their cultural context.
My intent: clarity. Their interpretation: threat.
Neither of us is wrong about what the words mean in our own dialect. The problem is the missing translation layer. There’s no shared understanding of what direct communication means in my context versus their context.
The manifesto was my first attempt at building that translation layer for humans:
“We’re direct and sometimes unfiltered. If something lands badly, say so. We own impact and correct fast.”
That’s a human protocol for what the corpus does with data — acknowledge that the same signal can be interpreted differently in different contexts, create a feedback loop to correct mismatches, build a shared understanding over time.
“Is it me, or is it them?”
That’s the golden record’s core question. When two sources disagree, which one is authoritative? Sometimes neither. Sometimes both. The answer is: understand both contexts, build the canonical version, let field authority determine truth based on quality, not hierarchy.
Multiple jobs. Same pattern.
Every company I worked in had internal dialects. Sales spoke one language. Engineering spoke another. Leadership spoke a third. Each department adapted business concepts to their context. Each one was internally consistent. None of them aligned with each other.
I pointed this out. In every company. For 18 years. The response was consistent: “That’s just how it works.” “Stay in your lane.” “You’re a salesman, not an engineer.” The translation layer — the person who could see all the dialects and identify the misalignments — was treated as the problem rather than the solution.
Multiple jobs. Every exit came from pointing out that the internal dialects weren’t aligned and nobody had built a shared understanding layer.
The ADHD accelerated the pattern recognition. I saw the misalignment faster than most people. I named it faster. I couldn’t not name it. That’s why I hit targets in 15-20 hours — I saw the shortcut through the dialect maze that others were still navigating. And that’s why I got fired — I told people their maze existed.
The platform is the translation layer I couldn’t be
I couldn’t be the lingua franca for every organisation I worked in. A human can’t be a translation layer for a company — it’s too much context, too many dialects, too much political complexity. The companies didn’t want the misalignment pointed out. They wanted the misalignment to be invisible.
So I built a system that does what I couldn’t do as a person. It sits between dialects and creates shared understanding. It doesn’t judge any dialect as wrong. It doesn’t tell Salesforce to rename their fields. It doesn’t tell the sales team to speak like engineers. It understands both, maps the overlap, preserves the differences that matter, and creates a canonical version that everyone can work from.
The manifesto is the translation layer for people. The manifests are the translation layer for code. The corpus is the translation layer for data. All three solve the same problem at different scales. All three came from the same experience: 18 years of being the person who saw the misalignment and wasn’t allowed to fix it.
VIII. The Endgame — Lingua Francas Don’t Get Replaced
Lingua francas do not get replaced by the next contender. Once enough of the world speaks one, it only becomes more entrenched.
Once established, they persist
Latin lasted as the lingua franca of Western Europe for over 1,500 years. Even after the Roman Empire fell, Latin persisted in law, science, religion, and diplomacy for centuries. It was never “replaced” — it evolved into French, Spanish, Italian, Portuguese, and Romanian, each of which carried Latin’s structural DNA forward.
English has been the global lingua franca for roughly 400 years and shows no signs of being displaced. Not because English is “better” — because the network effects are too strong. Too many people speak it. Too much knowledge is encoded in it. Too much infrastructure is built on it. The cost of switching to a different lingua franca exceeds the cost of any inefficiency English creates.
Once a lingua franca reaches critical mass, it becomes self-sustaining. Adoption drives more adoption. Each new speaker makes the language more useful for every existing speaker. The cost of not speaking it increases with every person who does speak it.
The corpus follows the same trajectory
At 1,941 providers, the corpus is useful. At 30,000-45,000, it’s comprehensive. At 100,000+, it’s the default. And once it’s the default — once enough providers are classified, once enough customer data flows have validated the mappings, once enough partner platforms have embedded the SDK — the network effects make displacement functionally impossible.
A competitor starting today would need:
- 26 format parsers (months to build)
- 1,941+ providers ingested and classified (months to years)
- 234,373 vendors mapped with competitive relationships (years — G2 took years)
- Millions of validated field mappings from real customer data flows (impossible without deployment scale)
- And by the time they match today’s corpus, it would be 10× larger
The corpus is a lingua franca. Lingua francas, once established, don’t get replaced. They get adopted.
What this means for the software industry
The integration problem has existed since the first two software systems needed to communicate. Every solution so far — EDI, ESB, iPaaS, unified APIs — has been a partial bridge. A manual translation. A consultant-dependent, connection-specific, breakage-prone translation layer that treats each pair of systems as a unique problem.
The corpus treats it as what it is: a language problem. 234,373 dialects of the same underlying concepts. People. Money. Transactions. Products. Time. Relationships. Permissions. Events. The dialects are different. The semantics are universal.
The solution isn’t better connectors. It’s comprehension. Understand what every dialect means. Map the structural overlap. Preserve the meaningful differences. Create the canonical version. Absorb changes as dialect evolution, not breakages. Document the etymology. Keep the conversation flowing.
That’s what languages have done for millennia. That’s what the corpus does for data.
The Summary
Allow everyone to be different, but understand the differences and adapt yourself — don’t try to adapt them. It works both ways.
That principle, applied to matter, produced chemistry. Applied to organisms, produced ecosystems. Applied to communication, produced language. Applied to knowledge, produced science. Applied to commerce, produced trade. Applied to industry, produced standards. Applied to data, it ends integration.
The pattern is always the same. It has been the same since the Big Bang scattered matter across the universe and each clump adapted to its local conditions.
- Something new emerges (matter, life, language, religion, industry, software, AI)
- Each group adapts it to their context
- The adaptations become locally optimised and globally incompatible
- Communication between groups becomes necessary (trade, cooperation, compliance)
- Manual bridges appear (translators, consultants, middleware — expensive, slow, one pair at a time)
- A lingua franca emerges (Latin, English, TCP/IP, the Inferrex corpus)
- The lingua franca absorbs from every group, becoming richer than any individual dialect
- Critical mass makes adoption inevitable — the cost of not speaking it exceeds the cost of learning it
- The lingua franca becomes invisible infrastructure
- The next wave of creation starts the cycle again
Evolution doesn’t produce perfection. It produces adaptation. Religions adapted the same spiritual concepts to local cultures. Empires adapted governance to conquered territories. Industries adapted tools to local manufacturing needs. Software vendors adapted business concepts to their data models.
Nobody was wrong. Everyone was locally optimised. The incompatibility only matters when they need to communicate — which is always.
The software industry is in stage 5 — manual bridges (iPaaS, consultants, custom code) that cost six figures and break constantly. The Inferrex corpus is the transition from stage 5 to stage 6 — from manual bridges to a lingua franca. Once it reaches critical mass, it becomes stage 9 — invisible infrastructure that everyone uses without thinking about it.
And the AI wave that’s fragmenting everything right now? Same pattern. Stage 2 — every vendor adapting AI to their context. Different dialects of the same concepts. OpenAI’s “completions,” Anthropic’s “messages,” Google’s “candidates.” The corpus absorbs them all. Because the pattern doesn’t care what the medium is. Matter, language, religion, steam, electricity, software, AI — the cycle is the same. The lingua franca is the same solution. And it’s been the same solution for 13.8 billion years.
The Industrial Revolution needed standard gauge railways. The internet needed TCP/IP. Global commerce needed English. The software industry needs a lingua franca for business data.
1,941 providers. 234,373 vendors mapped. 38 formats normalised. 21.1 million integration pairs growing toward 1 billion. Self-healing that comprehends changes rather than fixing breakages. A temporal model that preserves every version like the OED preserves every word. A golden record that creates the canonical version without judging any dialect as wrong.
That’s not integration. That’s comprehension.
And the person who built it is someone who spent 18 years being misinterpreted, built a manifesto to fix it for people, built manifests to fix it for code, and built a corpus to fix it for data. The same solution at every scale. Because the problem has always been the same: different groups adapting the same concepts to their local context, and nobody building the shared understanding layer.
The lingua franca of business data. Day one. And lingua francas, once established, don’t get replaced.
The Big Bang fragmented matter. Evolution adapted life. Culture adapted belief. Industry adapted tools. Software adapted data models. AI is adapting intelligence. The pattern is 13.8 billion years old. The solution — comprehension, absorption, shared understanding — is as old as the first two organisms that needed to cooperate.
Inferrex isn’t a software company. It’s the latest instance of the oldest pattern in the universe: local adaptations becoming globally comprehensible through a shared understanding layer.
And this time, the bridge builds itself.
Allow everyone to be different, but understand the differences and adapt yourself — don’t try to adapt them. It works both ways.
That’s the golden record. That’s the manifesto. That’s the lingua franca. That’s the principle that’s been the right answer since the beginning of time — and the one the software industry finally has the tools to implement.
inferrex.com
As of 4 September 2026 · inferrex.com/claims
Aaron Gammon — Founder
Inferrex · Inferrex Ltd · Company 17011782 · London, UK

