Build eval harnesses and benchmarks that use tracked pricing outcomes as ground truth for model quality.
Systematize and automate expert review workflows that are currently handled manually.
Develop AI personas that simulate B2B buying committees using usage data and call transcripts.
Automate persona training pipelines that today require manual effort.
Own LLM routing across providers such as Anthropic and Google, with explicit cost, latency, and quality tradeoffs.
Maintain infrastructure and data residency boundaries, ensuring regional model calls stay within the correct geography.
Extend the MCP server used by LLM agents so that product features are agent-driven, not just UI-rendered.
Work within a typed ontology of pricing entities to keep model outputs structured and auditable.
Identify and remediate systemic latency, data drift, and cold-start issues in the pricing loop.
What We're Looking For
8 or more years of engineering experience with strong, recent production LLM depth.
Proven track record shipping and owning LLM-powered product features end-to-end, from development through production monitoring.
Direct experience building evals and observability for LLM systems, including creating eval harnesses and baselining prompts against typed ontologies for release gating.
Hands-on experience with MCP or building tools for LLM agents, and familiarity with platforms such as LangChain, LlamaIndex, Braintrust, or OpenRouter.
Experience with structured data models and typed schemas (for example, Pydantic) to ensure model outputs land in auditable form.
Experience operating under data residency, SOC2, and GDPR constraints in domains where correctness is audited, such as pricing, billing, or payments.
Strong communication skills to explain non-deterministic systems to clients and partners.
Product-engineer instincts to scope and ship pragmatic solutions in a lean, fast-moving environment.
Must be authorized to work in the United States; visa sponsorship is not available.