CERTAINCE Logo

The Best AI Model or Better Context for Office Work?

AI Content · 11 min

Illustration: a large brain built from geometric shapes beside a small, well-sorted filing cabinet with a golden line running to a desk

An employee has a reply to a complaint drafted for her. She receives a well-built text, friendly in tone, citing a goodwill rule the company does not have. A larger model with better exam scores would have written the same sentence, only more elegantly.

Everything was on the model's desk except that rule.

Since ChatGPT became public in November 2022, every model generation has improved at exams, at programming and at long chains of reasoning. In everyday office work that gain is harder to feel, because reasoning power is rarely the scarce part there. A project we built ourselves shows where a stronger model helps and where better context matters more.

What a better model actually changes at the desk

The progress is real, and it does show up at the desk. Current models can follow an instruction with several exceptions and keep the structure of a twenty-page document. They can also write a German quotation without someone having to rework every second sentence. Anyone who tried once in 2023 and gave up afterwards should try again; the tools of that time are no longer a yardstick. What improved is the handling of whatever is in front of the model. What a model knows about your company has not changed.

The limit of that progress can now be quantified. How well AI systems answer questions about weeks of work in the same environment is what a UCLA group tested in May 2026 under the name LongMemEval-V2. Frontier models that cannot look into that work reach 14.1 % correct answers – with the same reasoning ability that makes them shine in other exams. An off-the-shelf tool of the same kind, allowed to search the histories as ordered files, reaches 69.9 %.

The measurements ran in two artificial English test environments where neither permissions nor conflicting sources nor data protection play any role; even the best setup lands at 74.9 %, which means one answer in four is wrong. The full assessment is in context management for business. For model selection, this one result is enough: the gap between 14.1 % and 69.9 % comes from access to your own history, not from a smarter model.

A project where that difference became visible

MAFU-SHERPA CNC Automation builds robot cells that load turning and milling machines automatically. For selling those cells, a company record such as “machinery, 50 to 200 employees, Baden-Württemberg” is close to worthless. It says nothing about which CNC machines are on the shop floor, which parts are produced and how often the setup changes. None of that sits in an address database. It sits on the machine-park page of the company website, in job ads for CNC operators and in investment announcements.

We built this company a custom sales solution that reads exactly those sources and checks what it finds against the criteria for a fitting company, defined beforehand. If it finds a machining centre on the machine-park page, the draft argues flexible loading and additional spindle time. If it finds open positions for CNC operators, it speaks about relieving the existing shift. Every company gets a fit between 0 and 100 and a justification that names the machines and signals actually found. Roughly twenty work steps run one after another, each one costed separately. A member of the sales team reads the draft and sends it.

The models inside it are generally available ones. The difference comes from having the company's website and the marks of a good match in front of them. A rule from our own work: first someone writes down how a fitting company is recognised. Only then may a tool start reading. In the reverse order what comes out is a summary nobody in sales can use.

What a new colleague learns in her first weeks

Someone who starts a new job knows the profession and still keeps learning for weeks. She learns where the current prices are and which file is the old one. She learns that one particular customer always wants the delivery date in writing. She learns which promise cost money last year and has not been made since. None of this is in her employment contract, and most of it is written down nowhere at all. After six months she works considerably faster than on her first day, without having become any smarter.

A language model brings the professional groundwork with it, across a breadth no single person has. The second part it does not bring, and it does not gather that part by itself either.

It starts every day as its first day.

In-context learning: learning for the duration of one task

What makes a model adaptable regardless is called in-context learning. The model aligns its behaviour purely with what stands in the current request, that is with the rules, examples and prior cases someone supplied. Three genuine earlier replies usually work better than a page of style guidance; the technical term for this is few-shot. Nothing of it is learned permanently. After the task the model is in the same state as before, and the next request starts from zero again.

For most office tasks that is enough. A quotation email, the summary of a set of minutes or the preparation of a report are tasks a knowledgeable person also completes with the documents in front of her.

The MAFU-SHERPA sales solution learns nothing either. For every company the same criteria and the same product knowledge are supplied again. That it still runs dependably comes from both being written down and not sitting in one colleague's head.

How far that carries has been measured at a real workplace. Erik Brynjolfsson, Danielle Li and Lindsey Raymond evaluated the rollout of an AI assistant to 5,172 customer-support staff in the Quarterly Journal of Economics, with 15 % more issues resolved per hour on average. Less experienced staff became faster and better, while the most experienced ones barely gained speed. The effect was largest on rarer problems for which the company already had enough earlier cases. What the assistant passed on there was what the business already held – and those 15 % come from one customer-support setting and are no general office calculation.

Why more context is not automatically better context

The context window, meaning the amount of text a model can process at once for one request, has grown considerably in recent years. The obvious response to that – pour everything in – helps less than expected. In the same UCLA measurement, a simple automatic search across the histories reaches 42.8 %, while the same histories as ordered files produce 69.9 %. The single most effective component was a plain document describing the workflow: leave it out, and correct answers in the large run drop from 70.1 to 64.1 percent.

The same selection work sits inside the sales solution. What gets evaluated is whatever says something about machines, batch sizes and capacity; the rest of the website stays outside.

The hard part, therefore, is choosing the right material. Which source leads, who maintains it, who may see it and how quickly it goes stale is decided before the tool is chosen. That is context management, and it is the part of an introduction that no change of model takes off your hands.

You can require these decisions of a provider before you introduce a tool. Six questions cover them, and none of them assumes any software knowledge.

  • Which task exactly should the tool take over, and how will you know in everyday use that a result is right?
  • Which source does every statement in the result come from?
  • Who in the company maintains that source, and what happens when that person is away for two weeks?
  • Who may see which documents, and does the tool see more than the employee operating it?
  • How old may a piece of information be before it must no longer be used?
  • Can you trace, on a finished result, what it rests on?

Continual learning: what is missing today

A colleague delivers a third step as well: she keeps a correction. Explain once that a customer does not want collective invoices, and it does not have to be repeated in every email. That is exactly what today's models do not deliver. What is offered as a memory function is storage and retrieval. A note is saved and sent along again with the next request, while the model itself remains unchanged.

For permanent learning there is the term continual learning: a system that keeps what it has been corrected on during ongoing operation, without anyone supplying the same correction again. Whether and when that arrives as a dependable function in office tools is open, and it remains to be seen whether it arrives in the shape people expect of it.

For a company it would be only half the answer in any case. A system that learns unnoticed also learns the wrong things – an exception that applied only once, or a promise a colleague made in error. Such learning becomes usable only when what has been learned is visible and can be withdrawn. How to solve that technically goes beyond the scope of this article.

When a stronger model is the right answer

There are tasks where the choice of model makes the difference. You recognise them by the fact that all the necessary information is available and the task stays hard anyway.

  • A long contract in which two clauses contradict each other and someone has to name the consequence.
  • An analysis meant to derive a relationship from the available figures that nobody has pointed out.
  • Programming tasks across several files where a mistake quietly produces wrong results instead of standing out.
  • Texts in a language nobody in the company can proofread.

The rule of thumb in between is simple. If an experienced person could complete the task with exactly the documents the model received, and the result is still wrong, a stronger model is worth it. If that same person would have to ask around first, context is missing, and a stronger model only guesses with more confidence.

A test that costs one afternoon

Before you decide about tools or plans, take real cases from last week. In our projects we work with three to ten cases, among them at least one entirely ordinary one, one exception and one that went wrong. That is the second rule from our work: a tool checked only against the easy cases fails later on exactly the cases that make the task worth doing.

For each case, put together everything an experienced person would have needed for it, and give that to the tool you already have. If the result is right now, the task is a question of supply; someone has to make sure those documents are available every time, without a person collecting them. If it is still wrong, you also know more, and the question about the model is justified.

The test costs an afternoon and no licence. It is the same first step we recommend for every task where using AI in your business is meant to start.

Does my company need the best AI model for office work?

For most office tasks, no. Between a very good model and the current best one the difference barely matters there, because the result depends on the information available. A stronger model pays off with long contradictory documents, with programming tasks and wherever all the documents are available and the task stays hard anyway.

What is in-context learning?

In-context learning describes the ability of a language model to derive its behaviour purely from what stands in the current request, that is from the rules, examples and prior cases supplied with it. Three fitting examples often work better than a page of guidance. Nothing is learned permanently: after the task the model is in the same state as before.

What is continual learning?

Continual learning describes a system that keeps learning permanently during ongoing operation, meaning it retains a correction it was given once without being told again. Today's memory functions do not deliver that; they store notes and present them again with the next request. Whether and when dependable learning arrives in office tools is open.

Why does the AI forget what we discussed last week?

Because every request stands on its own. The model retains nothing by itself; available is only what is supplied for that one request. What looks like memory is a stored note the tool sends along again. It becomes dependable only once the information sits in a maintained source and not in a chat history.

If you have a task in mind where an AI tool comes out almost right today, a conversation can narrow down what those last few percent depend on. Book an intro call.

Inquiries