Back to careers

Research agents, retrieval and evaluation

Agent Engineer

The mission: Build agents that research, reason with evidence and use tools to deliver answers people can inspect and act on.

We hire according to current project needs. We welcome expressions of interest in these roles; timing and scope depend on our priorities.

Where you’d apply it today

Choose

Sourced buying research: turn a buying brief into a comparison people can inspect, under real constraints and explicit uncertainty.

What you’ll do

  • Design the agent harness: planning, search, source inspection, clarification and structured results, with clear task constraints and stopping conditions.
  • Engineer context across long tasks: retrieve evidence when needed, preserve source references, and use memory and compaction without losing constraints or unresolved questions.
  • Build tools with clear schemas, useful error responses and validated outputs; support durable checkpoints, cancellation and safe retries.
  • Turn real failures into evaluation datasets. Assess grounding, tool choices and task outcomes with automated checks and human review; compare model and prompt changes against quality, latency and cost.
  • Keep sensitive information protected, treat retrieved content as untrusted, and communicate uncertainty or an honest no-fit result.

What you’ll bring

  • Strong Python or TypeScript engineering, including async APIs, tests and debugging. Be ready to work in our TypeScript/Node.js stack.
  • Hands-on LLM application work: tool calling, structured outputs, prompt design and context management. Explain when a simple workflow is sufficient and when an agent loop helps.
  • Practical retrieval and search skills: selecting sources, preserving citations, handling conflicting evidence and choosing what enters the model’s context.
  • Ability to design evaluations, read execution traces and improve an agent based on measured failures rather than a convincing demo.
  • Product judgment: translate a vague request into testable constraints and explain trade-offs clearly to users and teammates.

Helpful experience

  • Agent SDKs or orchestration libraries, such as the Claude Agent SDK, LangGraph or Vercel AI SDK; experience with tool integrations through MCP (Model Context Protocol).
  • Hybrid search, reranking, multimodal evidence or long-task memory; experience deciding when these techniques improve results.
  • React/Next.js, PostgreSQL, experiment tracking, model routing or prompt caching. Experience with AI coding tools and reviewing their changes critically.

Why you’ll find it rewarding

  • Work on the whole path from an ambiguous request to a useful, evidence-backed result.
  • Shape the prompts, tools and evaluations together, and see which changes actually improve the agent.
  • Build capabilities that can travel across products, starting with research and decision support.

How to apply

Send a short introduction and links to relevant code, projects or a portfolio; a CV is optional. Tell us what you built, a failure you investigated, and how you checked whether it worked. We value demonstrated ability over particular degrees or years of experience.

Apply by emailOr write to us directly: contact@dpintelli.com

Opens your email app with a prepared draft addressed to us.