Location: San Francisco (Hybrid: 3 days in SoMA) | Open to London (5 days/month in SF)
Type: Full-Time
Team: Research
Focus: AI & Agent Systems
Who We Are
We're a team of founders, engineers, researchers, creatives, and operators building what we believe will be one of the defining companies of the AI era.
Some of us came from startups. Some from academia. Some from hospitality, defense, design, or completely different worlds. We don’t care how you got here - we care about your energy, intelligence, curiosity, resilience, and ambition.
Who We Are Not
- We are not your average SF startup. We’re a diverse team of 20 people - hailing from 10 countries and speaking 10 languages with a range of non-traditional backgrounds.
- We are not a remote-first zoom culture: We value the energy of solving hard problems on a whiteboard in our SF HQ.
- We are not “researching to learn”: Research and product move in tight, high-velocity loops. If it doesn’t become code to solve a customer problem, it doesn't exist.
What You’ll Own & Build
As a Member of Technical Staff within the Research Tribe, you’ll be one of the early engineers shaping the core systems that power Seldon.
You won’t just build agents - you’ll design the infrastructure that governs how they operate: how they coordinate, how they’re evaluated, and how their behavior improves over time.
This is a deeply technical role bridging distributed systems, machine intelligence, and product - turning stochastic models into reliable, production-grade systems.
- Multi-Agent Systems Architecture: Design and implement agentic systems, rather than prompt chains, at scale. You’ll define agent topologies, robust orchestration layers, structured state management, persistent and scoped memory, tool-use registries and execution, and build structured, inspectable systems. You’re building a long-running harness-based agentic platform, not wiring together LLM API calls.
- Behavioral & Persona Modeling: Simulate user archetypes with goal-directed behavior. You’ll design data-grounded personas, enable “computer-use” on real interfaces, and develop agents with critique and reflection loops. You’re producing measurable, actionable behavior.
- Evaluation & Prompt Engineering: Design prompt schemas and reasoning templates as integrated systems. You’ll build frameworks that measure behavioral fidelity and failure modes, and implement automated and HITL feedback. If it can’t be measured, it doesn’t ship.
- Signal & Insight Systems: Build systems that transform raw agent outputs into high-signal insights. You’ll cluster observations across runs, implement severity scoring, and map behavioral signals to product decisions. Turn noisy outputs into decisions.
- Production Reliability: Build the infrastructure to operate agents in production. You’ll implement tracing, attribution, and cost monitoring, detect failure modes and design fallbacks, circuit breakers, and evaluation pipelines. The bar is production reliability.