← All insights

Artificial intelligence

Why your AI agents get your data wrong

22 September 2026·6 min read

An AI agent always answers with confidence, even when it is off target. More often than not, the error does not come from the model but from a word that each department understands in its own way. An ontology gives the company a shared vocabulary.

An AI agent never says it hesitates. Asked for the month's revenue, it gives a figure, precise and delivered with assurance. That figure is correct under one definition. Nothing guarantees it is the finance director's.

The error almost never lies in the model. It lies in the data the company entrusts to it, and more precisely in the words that describe that data.

One word, several definitions

Take revenue as an example. Accounting reads it at the invoice date, excluding tax. Sales reads it when the order is taken. Operations, in a business of sites or projects, reads it by percentage of completion. Treasury only recognises it when cash is received. None of these readings is wrong, each answers a need. The word “customer” suffers the same fate: a legal entity for finance, an account for sales, a delivery site for logistics.

As long as people bridge the gap between departments, these differences are settled with a conversation, or a “oh, you mean the invoiced amount”. An AI agent does not talk to the next department. It queries the tables it is pointed to and settles on one reading, unaware that others existed.

What an ontology changes

Thomas Gruber defined it, as early as 1993, as an explicit specification of a conceptualisation. Applied to a company, the idea is concrete: describe the concepts that structure the business (customer, contract, order, invoice, site, product), their properties, the relationships that link them and the rules that apply to them. An invoice is tied to an order, a customer belongs to a group, a down-payment invoice is not recognised revenue.

A data dictionary says what each column contains. An ontology goes further, because it makes the links and the rules explicit, in a form a machine can use. That is what allows an agent to know which revenue it is being asked about, and to say so.

An ontology does not stop at isolated definitions, either. It also captures the relationships between concepts: a site depends on a customer, relies on subcontractors, who themselves depend on suppliers. This relational structure allows an agent to answer an impact question, not only a definition question.

Diagram: a supplier linked to a subcontractor, then to two sites and their respective customers
An ontology also links concepts to one another: it makes it possible to trace from a supplier through to the customers affected.

From ontology to answer: the semantic layer

This structure is not enough on its own, though. It needs a mechanism that connects it to the company's real data, to its tables, its systems, the flows that feed daily operations. That is the role of the semantic layer: an intermediary that translates a question asked in plain language into the right definitions, the right tables and the right filters.

Asked for revenue by product for the current quarter in Europe, an agent backed by a semantic layer does not guess the answer. It consults the ontology to find out what “revenue” means in that context, retrieves the corresponding data, and can cite the definition it applied. Without this layer, the same question requires the agent to choose alone between several tables, several filters and several conventions, without ever saying so.

Three-step diagram: ontology, semantic layer, AI agent's answer

Harmonise before automating

A difference in definition goes unnoticed in a monthly report reviewed by specialists. It becomes a problem as soon as an agent answers dozens of questions a day, asked by people who do not check. Automation does not create the inconsistency, it spreads it.

Two situations are then possible. In the first, the agent relies on a shared and documented vocabulary, its answers cite the definition used and can be checked. In the second, it improvises from implicit definitions, and no one can say which one it applied. The difference does not come from the power of the model. It comes from the work done on the data before plugging it in.

E-invoicing offers a large-scale example. The European standard EN 16931 describes a semantic model of the invoice, with data defined and named the same way for all senders and all recipients. It is a form of ontology imposed from outside, and it is what makes automated exchange between companies possible. Every company benefits from making the same effort for its own concepts.

Where to start

There is no need to model the whole company. A dozen debated concepts are enough to get going, those whose figures differ from one department to another: revenue, margin, customer, order, site or project. For each one, it is advisable to appoint an owner, write a definition, name the reference source, then make this repository available to the agent.

This exercise holds a surprise. It quickly reveals disagreements the company was carrying without knowing it, and that no tool can settle on its behalf. Someone must decide what “customer” means. This is the longest part of the process, and it cannot be delegated to an agent.

References

  • – GRUBER Thomas R. (1993), “A translation approach to portable ontology specifications”, Knowledge Acquisition, vol. 5, no. 2.
  • – European standard EN 16931-1, Electronic invoicing, part 1: semantic data model of the core invoice.

A similar topic to explore?

Book a meeting