Back to careers
Judges, verifiers and evaluation systems
Research Engineer — Agent Evaluation
The mission: Make agent quality measurable: develop evaluators that can check evidence, assess completed work and explain when a result needs human review.
We hire according to current project needs. We welcome expressions of interest in these roles; timing and scope depend on our priorities.
Where you’d apply it today
Vowdo
Develop evaluation for enterprise workflows: check outputs against task requirements, verify observable changes and identify failures before results are relied on.
Choose
Evaluate source grounding, constraint compliance and comparison quality, including conflicting evidence and cases where no option fits.
What you’ll do
- Turn workflow requirements into benchmarks, grading rubrics and versioned datasets drawn from consented, sanitized examples and verified outcomes.
- Combine deterministic checks, model-based judges and human review. Verify observable task outcomes as well as the quality of an agent’s response.
- Calibrate graders against independent human labels; measure disagreement, false approvals, uncertainty and performance across task types and languages.
- Build specialist evaluators through prompting, fine-tuning or distillation when experiments demonstrate an improvement over simpler baselines.
- Investigate evaluator bias, prompt injection, reward hacking and evaluation leakage; keep held-out test data separate from development and training.
- Integrate evaluation into regression checks and release decisions, with reproducible runs and clear quality, latency and cost reports. Model scores inform review; runtime permissions remain separately enforced.
What you’ll bring
- Strong Python engineering, reproducible experiments and practical statistical analysis; communicate uncertainty and distinguish meaningful improvements from noise.
- Hands-on LLM evaluation: dataset design, grading rubrics, human-label calibration and analysis of task execution traces.
- Experience with PyTorch or a comparable training framework, model inference and fine-tuning or distillation of existing models.
- Ability to distinguish a checkable fact or state change from a subjective judgment, and choose an appropriate verification method.
- Careful handling of sensitive enterprise data, annotation quality, dataset versioning and train/test separation.
Helpful experience
- Preference datasets, reward models or process-level verification; multilingual and adversarial evaluation.
- Parameter-efficient fine-tuning, evaluator calibration or serving smaller specialist models.
- Evaluation pipelines, observability tools, active learning or human-review interfaces.
Why you’ll find it rewarding
- Define what successful agent work means and build evidence that the product meets that standard.
- Connect model research to concrete workflow failures and decisions about what to ship.
- Build evaluation capabilities that transfer across enterprise automation and buying research.
How to apply
Send a short introduction and links to relevant code, projects or a portfolio; a CV is optional. Tell us what you built, a failure you investigated, and how you checked whether it worked. We value demonstrated ability over particular degrees or years of experience.
Opens your email app with a prepared draft addressed to us.