Teams already running agents in production are the buyers of evals, observability, and durable execution. This map splits them into two threads: companies already building on the open-source stack, and companies whose product is the agent itself.
Tier A accounts have engineers who already name LangGraph or LangSmith in their job postings — the conversation is expansion, not introduction. Tier B accounts sell agents as their product, where eval and audit-trail pressure is heaviest. The exclusions section lists what was cut and why, because the cuts carry as much judgment as the picks.
All three surfaced through job postings that mention LangGraph or LangSmith by name — the warmest signal available, and one that turns discovery from pitching into confirming.
Hiring a Senior Backend Engineer (AI) in New York with LangGraph and LangSmith named in the posting.
Customer-facing agents inside regulated banking workflows need audit trails, eval coverage before release, and tracing when a conversation goes sideways. Building that in-house competes with shipping product.
VP Engineering · Head of AI · Platform lead — ready for a Lusha enrichment pass
"Your backend AI posting names LangGraph and LangSmith — as the agent surface grows across bank clients, curious whether eval coverage is keeping pace with deployment."
Hi Daniel,
Your Senior Backend Engineer (AI) posting names LangGraph and LangSmith, which usually means agents are moving from prototype to production.
For teams shipping conversational AI into banks, the next gap is eval coverage and audit trails. Klarna's support agent serves 85M users on this stack.
Worth comparing notes on how you gate agent releases today?
Best,
Izzy
Direct line first, switchboard fallback. Voicemail references Email 1 by subject line.
No note. The profile does the selling.
Hi Daniel,
One more angle on this. Teams that build evals in-house usually staff it a sprint at a time. The real cost shows up later: senior engineers maintaining a test harness every time a model version bumps, instead of shipping product.
That trade is what the managed eval layer removes. It also carries the SOC 2 posture your bank clients ask about in vendor reviews.
If release gating is manual today, 15 minutes would tell you whether this is worth anything.
Best,
Izzy
Second pass. If Email 2 drew an open, reference it in the first sentence.
Daniel, I sent a couple notes on eval coverage. No pitch here: the Klarna write-up on gating agent releases is worth the read even if we never talk. If it sparks anything, you know where to find me.
Hi Daniel,
Last note on this angle. If eval coverage isn't the pressing thing right now, there's a separate conversation about what your bank clients' risk teams ask for. I'll save it for another week.
If it's wrong timing or wrong person, one letter back works: (a) later, (b) who.
Best,
Izzy
Role-level: Lusha holds no verified email for a product or delivery leader at Posh, so this persona is sourced via LinkedIn. The messaging shift below is the point: same product, different value language.
Hi [name],
When a credit union's risk team reviews Posh's agent, the questions are always the same: what changed, what was tested, where's the audit trail.
89% of teams running production agents now treat observability as standard practice, largely because that answer closes reviews faster.
If risk reviews are adding weeks between "agent improved" and "agent live," worth comparing notes?
Best,
Izzy
Reason-why opener, business flavor: every conversational AI vendor selling into banks is hitting longer risk reviews this year, and I work with the teams shortening them.
No note.
Hi [name],
The other side of the same coin: when resolution quality dips after an agent update, your client feels it before your dashboard does.
Klarna gates releases against eval suites and cut resolution time 80% while expanding what the agent handles. Quality proof like that is starting to show up in renewal conversations.
Happy to show what that looks like for a team selling into banks.
Best,
Izzy
Second pass on the direct line.
[Name], I sent two notes on how agent teams prove quality to bank risk reviewers. Sharing because the territory is useful whether or not we talk: the vendor review checklist banks run on AI vendors is changing fast. Happy to send what I'm seeing.
Hi [name],
Closing this thread out. If quality proof isn't the live issue this quarter, the engineering side of it may be, and that conversation belongs with your engineering lead.
One letter back works: (a) later, (b) talk to engineering.
Best,
Izzy
"Hi Daniel, this is Izzy calling from LangChain. The specific reason I'm calling you: your team is hiring backend engineers who work in LangGraph, and when a team hits that stage, the next question is usually how agent releases get gated before a bank client asks. Do you have 90 seconds?"
"Most teams your size gate with spot checks until the first client incident sets the policy for them. So the straight question: when an agent change ships today, what tells you it didn't regress?"
Listen. Then: "Worth 20 minutes with your team to compare that against what the eval layer automates?"
I did, and that's why I'm calling. Thirty seconds and I'll earn the reply or leave you alone. Land or no?
Fair, most teams I call have something working. One question and I'm gone: when an agent change ships, what tells you it didn't regress? If that answer is solid, you genuinely don't need me.
Respect, the serious teams all start there. The question that ages badly is maintenance: who owns the harness next quarter when the model version bumps? Teams move when their best engineers become the test infrastructure.
The OSS is the point, you have done the hard part. The platform is what OSS cannot be: hosted traces, shared eval suites, and the SOC 2 your bank clients' vendor reviews ask for. It earns its keep the day a risk team shows up.
Understood, one calibration question: is that because releases gate cleanly today, or because nothing has broken publicly yet? If the second, the priority usually gets set by an incident. Cheaper to set it yourselves.
Surfaced twice in sourcing: an agentic fraud-prevention product line, and an AI Engineer posting that names LangGraph. Double-verified.
Fraud agents make block-or-allow decisions where a silent failure is a lost customer or a passed fraudster. Every decision needs a trace; every model change needs an eval gate.
VP AI/ML · Platform Engineering lead · Head of Fraud Product
"Your AI Engineer req already names LangGraph — as agentic fraud decisions scale, the question becomes whether every team traces and evals the same way, or their own way."
Senior Software Engineer, AI Platform posting names both LangGraph and LangSmith.
Compensation math is revenue-critical: an agent error lands in someone's paycheck. Eval coverage and regression gates before release are the difference between a feature and an incident.
VP Engineering · AI Platform lead
"When the agent's output feeds a paycheck, 'mostly right' is an incident report — curious how your AI Platform team gates releases today."
Research, compliance, and financial-crime agents produce work a professional signs their name to. Provenance and evaluation are not features here — they are the product's license to operate.
Analyst-grade output for investment banks has to be traceable to sources. At Series D scale, multi-team agent development without shared evals and tracing turns every release into a risk review.
CTO · Head of AI Engineering
"When a bank asks how the model reached that number, the answer can't be a shrug — how does the team trace an agent's reasoning today?"
Embedding law into agents means every output is a compliance judgment. Provenance, eval coverage, and regression detection are what let a regulated buyer trust the product.
"When the agent's output is itself a compliance judgment, evals stop being an engineering nicety — curious how the team measures agent quality release over release."
Multi-step document workflows fail quietly at depth: step twelve of a diligence chain goes wrong and the summary still reads clean. Tracing across long agent runs is the fix, and building it in-house is a permanent tax.
CTO · Head of AI · Platform Engineering lead
"Deep document chains fail at step twelve, not step one — how does the team see inside a forty-step agent run today?"
Fraud and AML agents operate under examiner scrutiny: every automated decision may need to be reconstructed months later. Durable execution and full traces are audit requirements wearing engineering clothes.
CTO · VP Engineering · Head of AI
"When an examiner asks why the system blocked that transaction in March, the trace is the answer — how long do agent decisions stay reconstructable today?"
Agents are embedded across a very large engineering org, which usually means several teams solving observability separately. The enterprise conversation is standardization: one tracing and eval layer instead of five.
Head of AI · Platform Engineering leadership · Engineering directors per product line
"With agents shipping from multiple teams, the expensive question isn't whether to trace — it's whether everyone traces the same way."
Klarna — now a flagship customer (85M users served, 80% resolution-time reduction). Moved from prospect column to proof-point column.
Sentry — builds its own observability; a build-not-buy culture for exactly this category.
Fiddler AI — competitor in AI observability, not a prospect.
Ten staffing agencies — filtered from the job-posting data; a consultancy naming LangGraph is reselling talent, not buying tooling.
Pinegap · Henry AI · IANS · Nitrogen · GitLab · Harness — qualified in sourcing, held for the second sequence cycle.
Confirmed stack, regulated buyer, New York. Shortest distance from first sentence to a real conversation.
Double-verified signal and the clearest cost-of-failure story in the set.
Completes the Tier A wave while the expansion messaging is warm.
Rogo, Norm Ai, Hebbia, Sardine, Ramp — heavier education, richer pain. Sequenced after Tier A conversations sharpen the talk track.
Sourced in tandem: three Exa Websets searches (research-agent companies; job postings naming LangGraph or LangSmith; enterprises shipping agent features) produced 97 raw candidates. Manual scoring handled dedup, fit logic, agency filtering, and the demotions above. Committees were verified through a Lusha enrichment pass (August 2026); the map shows names, titles, and verification badges — the contact file itself stays private.