Forward Deployed · AI Customer Onboarding
Deployment Agent
A first-pass deployment agent that turns a customer's messy transcripts, Slack threads, and policy memos into a spend-platform configuration, plus an audit log that says exactly what a rep can trust.
Getting a company live on a spend-management platform means translating call transcripts, policy memos, Slack threads, and rosters into a precise configuration by hand. It is slow, and a confident mistake in a spend control is worse than an open question.
A Claude Code skill plus Python checkers that turn a customer folder into a schema-valid platform config and an evidence-backed audit log, with coverage, citation, and cross-customer checks and a cold-rebuild harness.
When a company adopts a spend-management platform, someone has to turn their reality (discovery calls, policy memos, Slack exports, rosters) into departments, users, card programs, spend limits, and approval rules. Today a deployment rep does that by hand. I built the agent that does the first pass for an AI deployment company: a Claude Code skill plus small Python checkers that read a customer folder, research the platform's current product and API behavior, and produce two files. The config is the proposed setup. The audit log is every assumption, customer question, conflict between the customer's own documents, and request the API cannot fulfill, each with evidence and a workaround. I completed all five customer scenarios, including both stretch cases.
- 1.Config From Messy Sources: Handle a clean transcript and roster, a Spanish interview with a mixed Portuguese and English policy memo, a healthcare company with compliance asks the platform may not support, a contradictory Slack export, and a 150-rep temporary card program with a role-by-merchant-category matrix.
- 2.Audit Log as Half the Deliverable: Record every assumption, missing detail, internal conflict, and unsupported API request with quoted evidence and a dated source, so the rep knows how far to trust each line of the config.
- 3.Never Invent Data: Only one customer supplied usable employee emails, so the other four ship complete spend designs with an empty user list. Waiting for a real roster is safer than generating plausible addresses.
- 4.Submit-Ready vs Apply-Ready: Every config is ready for review but explicitly marked not ready to apply until named customer confirmations come back. The agent prepares files only. It never calls the platform or touches an account.
- ◆Skill, Not App: The pipeline is an instruction skill and small scripts that any fresh Claude Code or Codex session can follow end to end. No UI, database, or live API connection.
- ◆Schema-Valid Is Not Enough: A checker suite goes past schema validation: semantic and referential config checks, source coverage, citation verification against the quoted file, cross-customer claims, and whether every numbered clause in a customer's own document gets addressed.
- ◆Test the Tests: 105 eval assertions, 17 of which require a checker to stay silent. A regression script reintroduces past defects, and a recheck script damages inputs to prove each checker can actually fail.
- ◆Cold Reproduction: Fresh agents given only the skill path re-ran all five customers from raw sources while the original outputs were fingerprinted and left untouched. Two rebuilt to an exact structural match.
A deployment rep will trust the config exactly as far as the audit log is honest. So the agent is built to flag rather than guess. When a healthcare customer asks for a control the platform does not offer, the audit log names the gap, cites the dated documentation it checked, and proposes a workaround. When a Slack thread contradicts the meeting notes, the conflict is logged with both quotes instead of silently picking one. The config is the proposal. The audit log is what makes it usable.
The first cold run by a fresh agent exposed instructions the skill was missing. After fixing the skill, the second cold run found the customer's precedence rule unaided. That loop, run cold, find what breaks, fix the instructions rather than the output, is what turns one good result into a deployment capability that works on the next customer you have not seen yet.
The real product is the audit log. A rep needs to know which parts of the config are safe, which need the customer, and which the platform cannot do. I designed every check around that question rather than around making the JSON look complete.
All five customer scenarios completed, including both stretch cases. Every config validates with zero errors, 105 eval assertions pass, and cold rebuilds from raw sources reproduced the original structure.
