Back to DocuFlow

What every model call cost, and why

A platform that reads documents with models has to answer three questions at any moment: what did this document cost, which model said what, and where did the time go. DocuFlow answers all three from one ledger.

Spend by day, model,
queue and document

Every call is filed against the work that caused it, whether that was extraction, a second reader, a rule asking a question or a reviewer asking the document. Open a trace to read it call by call: model, tokens, cost, latency and why it stopped.

AI cost

24 hours7 days30 days90 days

Every model call the workspace made: what it cost, how long it took, and what it was for.

Spend

$163.83

last 30d

Model calls

82,565

41 failed

Median call

1.9 s

p95 8.4 s

Tokens in

611.2M

402.9M cached

Tokens out

18.4M

2.1M thinking

Spend per day

Cost of every model call, by the day it was made.

By model

Which model the money went to, and how fast it answers.

ModelCallsMedianCost
gemini-3.1-pro-preview18,4426.1 s$132
gemini-3.8-flash41,2081.9 s$29.80
gpt-oss-120b22,9150.9 s$2.41

By kind of work

What the models were asked to do.

Extract aetna-pa-whitfield.pdf

Extraction · pipeline

Cost

$0.0884

Took

16.3 s

Calls

4

Tokens

38.1k / 2.2k

Document aetna-pa-whitfield.pdf
Queue Prior authorizations
Statusok

4 calls, in order

  • Classify pages1.2 s$0.0021
  • Extract fields · page 16.8 s$0.0412
  • Extract fields · page 25.9 s$0.0388
  • Second reader · 4 low-confidence fields2.4 s$0.0063
    Modelgemini-3.8-flashvertex
    Tokens6,120 in · 214 out · 5,880 cached
    Finished becausestop

    Prompt and answer text is not recorded on this deployment.

Counted where
the data lives

Analytics and cost are grouped in the database, so their cost grows with the number of days and queues, not with the number of documents. The headline figure and the breakdown read the same query, so they cannot disagree.

Analyticslast 14 days

Auto-approved

76%

3,412 of 4,488 approved

Median processing

14.2 s

p95 41.8 s

Documents per day

By final status

approvedneeds reviewrejectedfailed

Included

Prompts off by defaultdocument text is only recorded when you switch it on
Separate permissiontotals and per-call traces are different grants
90-day retentiontraces are pruned nightly
Never in the wayrecording never throws and never slows a document

Bring the forms your team still retypes

A working session on your own documents: one queue, your fields, your rules, and a count of what would have gone straight through.