lamapixel
0%
Automatizace · 5 min read

What exactly goes to the language model and what does not

Anonymisation before sending, your own gateway to the models, and a scripted branch for sensitive queries. Seven questions to check it with any supplier.

A metal kitchen sieve with a black handle on a light background

Proposals for AI automation promise that the system will read an incoming message, pull out what matters and draft a reply. What they almost never finish is the second half of that sentence: where the message goes while being read, who owns the server it lands on, and how long somebody keeps it there.

And yet a client can answer that question themselves in an afternoon, provided they know what to ask. This text describes where text can be stopped before it is sent, and ends with seven questions that work on any proposal, not only ours.

One boundary up front: what follows is technique, not a legal assessment. Whether you may carry out a particular processing of data is for a lawyer to say, not for an automation supplier.

The model does not need to see everything in the message

The concept missing most often from proposals is anonymisation before sending. It means there is a step between your system and the model that replaces sensitive data with placeholders before the message leaves your infrastructure. The model receives structure and meaning, not identity.

On one project in the summer of 2026 this was a requirement from the outset. We formulated it internally on 31 May 2026 and described it to the client on 17 June 2026: a custom anonymiser runs between the system and the model, replacing sensitive data before sending, with separate infrastructure alongside it.

The practical point is what does not need to be sent. A model drafting a document needs the document type, the relationships between the parties, deadlines and amounts. It does not need a name, an identification number, an address or an account number - those are filled into the finished document on your side, after the text comes back.

Replacement has one condition that everything breaks on: it has to be two-way and consistent. When two different people become one placeholder in the same document, the model returns nonsense that looks fine.

The best way not to send something is not to send it at all

Anonymisation handles the data the model has to see. That leaves the second group: queries where the model adds nothing and adds risk.

On our own bot, which tells clients the status of their case, this has been settled since 10 March 2026: a query about a specific case is served by a scripted branch, not by the model. The model may see the case name and record headings; nothing that identifies a person. The bot did not stop being useful - general questions, phrasing and summaries are still handled by the model; it simply never receives the one part where a mistake would cost the most.

That decision is made once, on paper, and it is cheap. If it is not made, it makes itself: everything flows into the model, because that is the simpler thing to implement.

A gateway instead of a direct connection to one provider

The third place where the fate of your data is decided is the connection architecture.

On the project mentioned above, the core does not talk to model providers directly but through a single gateway that speaks several providers' formats. Decided on 12 June 2026, and the reason was said out loud: not to depend on one provider. An accompanying decision from 9 June 2026: no plug-and-play boxed solutions; the system switches between models by task type.

Two things follow for the client, and both only show up a year later. When a provider raises prices, changes terms or withdraws a model, what changes is configuration, not the system. And when it turns out a particular task must not leave a particular territory or a particular provider, that task can be rerouted on its own, without rebuilding the rest.

Whose account it is and where the server stands

The most practical question on the whole list, and one you can answer in a minute.

On the same project the client set up both the server and the model provider account themselves, in their own name, on 24 June 2026. That is not a formality: the account owner sees the call history, pays directly, can change the password at any moment, and on parting with a supplier has nothing to take over, because nothing ever left their name.

The opposite arrangement, where everything runs on the supplier's accounts, is not wrong in itself and is common on small automations. But it has to be deliberate, and what happens when the collaboration ends has to be written down. There is a separate text on who owns the scenarios and the accounts.

What the system is fed is the same question as what leaves it

The last layer concerns knowledge rather than operations. A system meant to answer according to your rules has to be filled with content. There are two routes here, and they differ in dependency.

Somebody else's closed database is fast and brings a tie to the content supplier: their price list, their terms, their decision about what stays in it. Your own know-how plus open public registers is a slower start and an asset that stays with you. On the summer 2026 project we chose the second route on 17 June 2026 and filled the system with the client's internal material.

The way it learns is connected to this. When a system has several sub-agents, one per task, each learns separately on its own data. One agent therefore does not know material that does not belong to its task, and that is the cheapest way to limit the blast radius of a mistake.

Seven questions to put to every supplier

Get the answers in writing, ideally in the proposal itself.

  1. Which specific model provider processes our messages, and where does it physically run?
  2. Which data from the message leaves, and which is replaced before sending?
  3. Is there a branch that never reaches the model at all, and what belongs in it?
  4. Whose account is it with the provider, and who pays for it?
  5. Who can see the history of queries and answers, how long is it kept, and who can delete it?
  6. Will our data be used for training? If not, what proves it - an account setting, a contract, or both?
  7. What happens if the provider raises prices or withdraws the model? How much work is it to switch?

The seventh question reveals the architecture more reliably than the first six. A supplier who answers "that is built in, it cannot be switched" has also answered the question of whose system it is.

Want a specific proposal reviewed

Send us the proposal on your desk, even if it came from somebody else. We will return a list of the places where it does not say what happens to the data, and the wording to use when asking for it.

Systems where the model works only with what it has to see are what we build as part of business process automation.

Write to info@lamapixel.com or call +420 775 599 009.

Need a hand?

Write to us and we'll figure it out together.

Book a consultation →