Implementation costs

What an AI assistant costs once the pilot ends

The cost of a production system depends on traffic, the length of the context, the share of cases passed to a person, and maintenance. Here is how to work it out for one clearly defined case.

Read the article

Estimated reading time: 6 minutes.

A breakdown of the running costs of an AI assistant

A pilot of an AI assistant usually ends with an invoice that worries nobody. A few weeks of work on selected cases, a narrow group of users, one or two sources of knowledge. That invoice, however, describes a different system from the one that has to carry your full traffic.

So the question to ask before deciding on an implementation is not "how much does a month cost". It is: how much does it cost to handle one case, and what happens to that figure when all of the traffic reaches the system, rather than a selected part of it. Below we set out what makes up that cost and how to calculate it. We give no specific rates, because model pricing changes several times a year while the structure of the bill stays the same.

What the bill is made of

The cost of an assistant breaks down into four items. Companies usually count the first one, while the last two decide whether the system stays in the company at all.

  • The model - the charge for the input processed and the answer generated. It depends on the length of the question, the number of document passages attached, and the length of the answer itself.
  • Infrastructure - the document index, hosting for integrations, queues, the event log and a test environment. This item is largely fixed and grows in steps rather than in proportion to the number of questions.
  • Supervision - the time of the person who checks answers, closes exceptions and reports errors. Usually the largest item at the start, and the one most often left out of the budget, because you pay it in your team's hours rather than on a supplier's invoice.
  • Maintenance - updating sources, correcting rules, and responding to changes in the systems the assistant connects to as well as changes on the model side.

This split has a practical consequence. Infrastructure and maintenance are spread across every case, so per case they get cheaper as traffic grows. The model and supervision grow along with the number of cases. As long as you keep them in a single budget line, you do not know whether more traffic will improve the bill or make it worse.

Test traffic is not a forecast

A pilot almost always works on the easier part of the distribution. It receives cases that are well described, complete and picked by someone who knows the system. Production gets whatever arrives.

Four differences move the bill the most:

  • Pilot users know how to phrase a question. The rest of the company asks imprecisely, so one case takes several exchanges instead of one.
  • In a pilot you connect one or two sources. In production more are added, and each one lengthens the context passed to the model.
  • You run tests during the working hours of the implementation team. Production traffic has peaks, which force you to keep headroom in the infrastructure.
  • Unusual cases only appear once the volume of requests grows. A pilot does not show them, and they are the ones that trigger the most expensive path.

The unit cost from a pilot is therefore best treated as a lower bound, not as a forecast. A sensible forecast needs a period in which the system receives traffic that nobody has filtered.

Context weighs more than volume

With an assistant that works on documents, the charge does not depend mainly on how many times somebody asked. It depends on how much content the system had to read in order to answer.

Three things drive the length of the context: the number of passages retrieved from the sources, the conversation history attached to later questions, and the system instruction, which with elaborate rules can grow to the size of a separate document. Adding a second source of knowledge does not raise the cost only for the questions that concern it. It changes the cost of every question, because the system first has to work out which source is the right one.

Hence the practical conclusion: measure the average length of the context per case, not just the number of cases. That is the figure that explains why the bill grows faster than the traffic.

The case the system did not close

Every case follows one of two paths. A case closed automatically costs the model usage plus its share of the fixed costs. A case passed to a person costs the same, and on top of that the time of somebody who has to read the history, reconstruct the context and answer. The second path can be more expensive than handling the same case without an assistant, because it adds the work of a system that settled nothing.

A wrong answer that made it through counts separately. Its cost is the correction, the contact with the customer, sometimes repeating the work. It does not appear on the invoice from the model supplier, but it burdens the same process. That is why the error rate belongs in the cost calculation, not only in the quality report.

The cost of one case is not the price of a request to the model. It is the weighted average of two paths: the case closed automatically, and the case somebody had to take over. Until you know the share of the second path, you do not know the cost.

The rule we apply when estimating the cost of running a system

Working out the unit cost

The calculation can be assembled from data you already have, or can gather within a single billing cycle.

  1. Define the case. One customer enquiry, one document read, one report. The unit has to be identical on both sides of the comparison.
  2. Measure the model usage for a case closed automatically. Take the average across the whole period, not across the best examples.
  3. Add the fixed costs divided by the number of cases in a month: infrastructure, licences and the test environment.
  4. Price the supervision. Multiply the minutes spent on checking and corrections by the hourly cost of the person doing it.
  5. Calculate the exception path separately: the model usage plus the full time of handling the case manually.
  6. Combine both into an average weighted by the share of each path. That share changes over time, so recalculate it every month until it settles.

Compare the result with the cost of handling the same case today, calculated in the same way and including your team's time. Without that second figure, the first one answers no question at all.

When to check the break-even point

Not in the first week. The sensible moment comes once the assistant has worked through a full cycle on traffic that nobody selected: with missing data, unusual questions and a peak in load. Any earlier calculation describes a sample, not a process.

Two thresholds are then worth checking:

  • The volume threshold - the number of cases at which the fixed costs are spread widely enough for the unit cost to fall below the cost of handling the case by hand. A process with a small number of cases may never reach it. That is not a reason to give up, but a reason to spread the same infrastructure across several processes.
  • The quality threshold - the share of cases closed without a person involved at which the time saved outweighs the cost of supervision. Below that share, the system adds work for the team instead of taking it away.

Both thresholds are recalculated after every significant change: a new source of knowledge, a change of model, a new integration or a clear rise in traffic. That is also the moment when you can see whether it is worth moving part of the work to a cheaper model and leaving the harder cases to a more expensive one.

Before the decision

Gather three numbers before you extend a pilot or sign a contract for an implementation: how many cases of this type the company handles in a month, how long one case takes today, and what share of cases the assistant closes without correction. The first two are in your own systems or in your team's schedule. The third can only be measured when the pilot works on traffic that nobody has filtered in advance. Once those three numbers are known, the decision to go ahead stops being a matter of judgement and becomes a comparison of two unit costs.

If you want to work this cost out for your own process, book a free consultation. During the conversation we agree on the unit of a case, the data sources and the way to measure them, and afterwards you receive by email the scope needed for a quote.

Free consultation

Contact

Work out the cost for your process

In the first conversation we agree on the unit of a case, the data sources and the way to measure them. Afterwards we set out what has to be measured before a quote for the implementation and its upkeep can be prepared.

  1. 01

    Process

    You describe a task that recurs in a predictable way and takes up your team's time.

  2. 02

    Unit and data

    We agree what counts as one case and which sources the system may draw on.

  3. 03

    Measurement

    We set out the data to collect during the pilot, so the unit cost can be compared with today's.

Leave your contact details

You can also write to: kontakt@futurefirst.pl