CERTAINCE Logo

Give Every New Task to an AI Agent Once

AI Automation · 8 min

Illustration: different office tasks passing through an AI agent into a reusable workflow with human approval

On 2 September 2026, three supplier invoices were waiting in CERTAINCE's Outlook inbox. The instruction to the AI agent was brief: process yesterday's invoices, store the documents in the correct folder, and move an email only after its invoice had been saved safely. The agent found the messages from Neon, Clockify and DigitalOcean, separated the PDFs from embedded logos, validated the files and then moved the three processed messages. Its final check of the exact inbox confirmed that no invoice email for the period remained.

The more valuable part came next.

We asked the agent to preserve the method for future invoices. That one request became two reusable workflows for Claude Code and Codex: "receive-invoices" for incoming supplier invoices and "create-invoices" for our outgoing invoices. Next time, an instruction such as "receive-invoices yesterday" is enough. The workflow already contains the rules for the date window, the exact mail folder, duplicate filenames, PDF validation and moving a message only after the document is safe.

This is why delegating should become a habit. Give every real task to an AI agent once before deciding that a person must keep it.

The first attempt is a measurement

An abstract question such as "Can AI take over our accounting?" is too broad to produce a useful answer. A real instruction reveals much more. The agent has to find the correct messages, distinguish attachments, reach the files, follow the rules and return a result that can be checked. You can then see which parts work and which do not.

The rule applies to the attempt, not to the approval.

Sometimes the agent completes the task. Sometimes it reaches a missing permission or prepares a draft that needs a decision. That is still useful because the boundary is now concrete. The missing piece might be API access, a business rule or a named person who can release the result. An observed workflow has replaced a guess about whether the task is suitable.

The first attempt may take longer than doing the work by hand. Access has to be arranged, unspoken rules have to be stated, and mistakes in the instruction have to be corrected. That is the cost of turning personal routine into a working method others can inspect. For a rare task, one-off assistance may therefore be the end of the story. For recurring work, we judge the saving on the second and third runs, once the rules already exist. We do not discard a useful workflow because setup took time, and we do not build a large automation around one successful attempt.

In both of our tasks, the specialist label initially hid the individual steps. "Process invoices" sounds like accounting. "Adjust the newsletter" sounds like marketing. The actual work involved reading, comparing, storing files, transferring text and checking results. The agent handled much of that while a person kept the decision about a payment, a price or a send.

Turn one instruction into a working method

The invoice run could have ended as a pleasant one-off saving. At the next invoice batch, however, we would have explained the mailbox again, how Berlin calendar days map to UTC, which attachments count as invoices and when a message may be moved. Instead, every correction became part of the reusable workflow.

The independent review added safeguards that were missing from the first instruction. The archive folder must actually sit under the correct inbox. Every page of a longer result list has to be read. An existing file must never be overwritten silently. An inclusive range ending on 5 September technically closes at the start of 6 September. Those rules now travel with every run while the user's request stays short.

That 2 September left us with a habit: after every successful first run, we decide explicitly whether to preserve the method for next time. The result may be a short instruction in a Markdown file, a saved agent workflow with fixed checks or, later, an integration with the existing systems. The task determines the right form. Often, a readable file is enough and new software is not yet needed.

The first run established the path. Every later one starts with rules that are already recorded.

A task may reach further than expected

A second request that day started with fetching the latest custom HTML template from Mailchimp. Through the Mailchimp API, the agent checked all 15 campaigns in the account and established which of the three user templates had actually been used. Mailchimp returned an empty ZIP file when the agent requested a templates export. The agent then recovered the complete HTML through the campaign-content endpoint.

It checked again that the other two templates were not referenced by any campaign and deactivated exactly those two. The follow-up task required both Outlook and Mailchimp. The approved German and English copy sat in an Outlook thread, while the design sat in Mailchimp. The agent created two new language templates and uploaded them to Mailchimp without changing the template already in use. Sending the campaign remained outside the task.

That example changed our view of what can be delegated. The work went beyond writing copy or operating one screen. It required reading vendor documentation, comparing account history, finding another route after an empty export and carrying approved material from Outlook into Mailchimp. A person might have treated these as several small errands. The agent could follow them as one task because the outcome, access and stopping point were clear.

Try every task without automating every task

Keep the first attempt broad and the permissions narrow. An AI agent can begin by reading, comparing and preparing drafts. Once money, dates, customers or commitments are involved, the workflow needs a clear release point. Decisions with serious consequences stay with a person.

Our delegation map places that boundary at each step of a process. It separates fixed rules, assistance, drafts with approval and bounded automation with sampling. The two methods fit together. The first attempt exposes the steps, while the map decides how far each step may go later.

A stopped attempt is useful as well. If the agent cannot reach the system, the next decision concerns access or an integration. If nobody in the company can define what "correct" means for the case, the task stays with a person for now. If the work happens too rarely to justify preserving the method, one-off assistance may already be the smallest useful solution.

Test the next ten tasks

Take the next ten real tasks that reach your desk. Include work that does not look like an obvious AI candidate. A supplier invoice, a campaign change, a research request, a recurring report or preparation for a meeting all count. Give the agent the complete task with four pieces of information:

  1. Describe the result that should exist at the end.
  2. Name the sources and systems it may use.
  3. Define which action needs a person's approval.
  4. State how a correct and complete result can be checked.

Classify each attempt as "complete", "partial" or "not suitable". After a complete or partial result, ask one more question: which instruction, check or connection would make the next run shorter? Preserve only what proved useful in the real task. When one of those tasks comes up again, compare how much of your own time it requires with the first run. If an existing tool handles the workflow reliably, keep it as the simplest solution. If the normal case can be stated as an unambiguous rule, use fixed automation; reserve an AI agent for steps where cases genuinely differ. Only then does the choice between a no-code platform and a code environment matter.

This exercise is the core of the AI enablement workshop in our setup stage. Over two days, your team hands whole tasks from its own work to Claude, ChatGPT or Copilot, reviews the results and learns to take a task back when the tool keeps missing the same point. Before that, we connect the programs your team already uses, so that a first attempt does not fail for lack of access.

The three invoices did not prove that accounting belongs to AI. They exposed one precise workflow that the agent handled in this run and whose rules are now ready for the next inbox. That is the reason for the habit: you learn what can be delegated by delegating it, and the second run starts where the first one ended.

Inquiries