Invoices, internal tickets, order exceptions, vendor files — the records that fall out of automation and wait in a queue for a person. An agent works them in your systems of record, and routes what it can't clear to whoever owns that call.
The rules engine handles the clean records, and the percentage it clears is the number in the business case. Everything else — the invoice whose PO is off by a line, the request that needs an approval nobody chased, the order with an address that won't parse — falls into an exception queue that somebody works by hand, in batches, at the end of the month. The automation didn't fail on the hard records. It never covered them.
The interesting question about a back-office agent is never what it does with the clean rows. It's what it does with row 341.
The run
Collection Pipeline
One isolated run per invoice in the queue, each with its own history, latency and cost. Rows 1–340 match their PO and receipt and post.
Row 341
SAP · Confluence
The PO is off by one line. It pulls the contract, finds the terms changed last quarter, and computes the corrected total in a sandboxed Python step rather than doing arithmetic in a prompt.
Row 341
Policy
Corrected variance is still outside tolerance, and the amount is over the approval threshold. It flags that one invoice and keeps clearing the rest — the batch does not stop.
Controller
Slack
Gets the invoice, the variance, the contract line it read from, and what it would do. Approves the corrected amount.
Row 341
SAP
Posts, with the approval, the approver and the full reasoning on the run's trace. The run had parked on the gate for two hours and resumed where it left off.
Nothing was worked in a batch on Friday, nothing was posted without a signature, and the record of why is attached to the record itself rather than to somebody's memory.
Finance Operations
Matches the PO, checks the terms against the contract, clears what reconciles and flags what doesn't.
IT & Employee Support
Access requests, provisioning, and the same question asked forty ways — with a gate on anything that grants permission.
Order & Fulfillment Ops
Order exceptions, addresses that won't parse, inventory that doesn't line up, returns and replacements.
Compliance & Vendor Review
Vendor files, document completeness, KYC checks — with a human signature where the rule says a human signs.
Different systems of record, same shape: a queue, a tolerance, an exception, and somebody who owns the call. Connect a system once and every agent in the organization can use it.
A new invoice in SAP or NetSuite, a ticket in Jira or ServiceNow, a row in a queue table, a form submission, a file dropped in Drive or S3. The history and the policy that govern this record are retrieved at the step, so it starts with the context a person would go looking for.
Match the PO, confirm the terms against the contract, query your Postgres or BigQuery for what happened last time. When it needs to compute rather than call, it runs Python in an isolated sandbox — arithmetic in a prompt is a guess with a decimal point.
A mismatch outside tolerance, a counterparty it's never seen, an amount over the ceiling you set. It flags that one record and keeps working the rest, and a malformed row fails on its own instead of taking the run down with it.
Not to a shared inbox. The owner gets the record, the mismatch and the reasoning on Slack, answers once, and it clears. If nobody answers, the run parks on the approval for hours and resumes when they do.
Before it can post
Per action you decide what it does alone: handle it, ask first, or do it and tell someone afterwards. Anything that moves money, grants access or writes to a system of record can be gated, and a gated run pauses with the full reasoning attached instead of proceeding optimistically. What it can't clear becomes a question with a named owner rather than a queue somebody works through on Friday.
Approval gates that pause the run, with the record, the variance and the reasoning attached to the request
Sandbox Mode until you say otherwise — including read-through, where reads hit the live system and writes are captured instead of sent
Every step traced: what it read, which tools it called, what it decided, how long it took, what it spent
Timeouts and a fallback ladder, so a record never sits waiting on one person being at their desk
Two of these are probably already in place and working. The question is what happens to the records they were never able to cover.
What that leaves
What this does instead
Deterministic, which is the point — and why the record that doesn't match the rule goes into a queue rather than through it. Every new exception shape is a change request.
It reads the contract, the history and the policy, corrects what it can defend, and escalates the rest with its working shown. The rules engine keeps the clean rows.
Very good at the one queue it was built for. The exceptions that cross systems — the invoice that needs the contract, the ticket that needs the entitlement — still land on a person.
One agent, your connected systems, and a directory of who decides what. The queue doesn't have to be a supported document type.
Headcount you can rent, at a cost per document that stops falling, with the audit trail sitting inside someone else's process.
The trail is yours, per record, exportable, with the approver named on it — and the cost per record is a run, not an hour.
The agent is a fortnight. The directory, the escalation ladder, the approval gates, the traces, the cost ceiling and the eval harness are the next two quarters — and they're the part that decides whether it ships.
All of that is the platform. Your team writes the agent and the prompts; nobody writes a retry queue.
Building it yourself is the one worth pricing properly — there's a page for that at /compare/diy.
Prebuilt tools for the systems your operation runs on — plus your own REST endpoints imported from an OpenAPI spec, and direct access to your databases when the catalogue doesn't cover it.
Slack
GitHub
Jira
Google Drive
Salesforce
MongoDB
Notion
Linear
Grafana
Discord
Google Chat
Confluence
PostgreSQL
HubSpot
Airtable
Shopify
Stripe
Datadog
Sentry
GCP
BigQuery
Slack
GitHub
Jira
Google Drive
Salesforce
MongoDB
Notion
Linear
Grafana
Discord
Google Chat
Confluence
PostgreSQL
HubSpot
Airtable
Shopify
Stripe
Datadog
Sentry
GCP
BigQuery
Google Calendar
Gmail
Mixpanel
Monday
MySQL
SAP
Zendesk
Zoho CRM
Google Maps
Google Ads
Coralogix
Telegram
Apollo
Mailchimp
Calendly
Redis
Supabase
GCP Logging
gVisor
Loops
Typeform
Google Calendar
Gmail
Mixpanel
Monday
MySQL
SAP
Zendesk
Zoho CRM
Google Maps
Google Ads
Coralogix
Telegram
Apollo
Mailchimp
Calendly
Redis
Supabase
GCP Logging
gVisor
Loops
Typeform
Don't see yours? Import tools from any MCP server.
The back office has a different risk profile than a support queue: mistakes are quiet, and they're discovered by an auditor. So the rollout is built around watching it work before it's allowed to write.
Week one
Read-through Sandbox Mode: reads hit the live system, writes are captured instead of sent. At the end of the week you have a list of what it would have posted, per record, with reasoning.
Weeks two and three
Turn on writes under a low threshold — the records your rules engine would have cleared anyway. Everything above the threshold routes to an owner. The comparison against last month is on one screen.
Month two
Move the thresholds per action rather than per agent. Most teams keep a permanent gate on anything that pays a new counterparty, and that's the right instinct — the gate costs one click.
Audit, control, and what happens when it's wrong.
What does our auditor actually see?
Per record: what the agent read, which systems it called, what it computed, what it decided, who approved it and when. Every reasoning step references a named, versioned agent rather than an anonymous prompt, every save is an immutable revision, and the trail exports.
Can it post to our ERP without a human?
Only for the actions and thresholds you configure. The common setup is automatic below a value you set, gated above it, and permanently gated for anything that pays a counterparty it hasn't seen before.
How does this sit with segregation of duties?
The agent is not an approver. It prepares and proposes; the approval gate routes to a named person by role, and the approver on the record is that person. Where your policy requires two, the ladder supports two.
What happens when it gets a record wrong?
The trace shows which step went wrong, and every change to the agent is an immutable revision you can roll back in one click. Changes can be gated behind evaluations that run against your own historical records before they deploy.
Does it do arithmetic in the model?
No. Computation runs as Python in an isolated sandbox and the result goes back into the run. It's the difference between a calculation you can re-run and a number that sounded right.
Someone's waiting
Support, success, onboarding — the work with a customer waiting on it.
Customer Support
Works the queue in your helpdesk, acts in the systems behind it, and asks before it invents a policy.
Customer Success & Renewals
Watches every account on the same schedule and hands the at-risk ones to the AE with the reasoning attached.