The hard problem in industrial AI isn’t generation. It’s judgment. Our agents draft and maintain FMEAs and score failure records for live enterprise customers, and every customer measures the system against their own best engineers.
We’re looking for a world-class ML engineer who wants to own that judgment: evaluation, calibration and the learning loop.
What you’ll do
- Build evaluation sets with customer engineers and measure accuracy where it matters most, on the highest-risk and lowest-risk calls.
- Calibrate scoring and reasoning across model providers, and swap models without losing accuracy.
- Turn engineer corrections into memory the agents reuse.
- Own quality, cost and latency for every deployment.
Who you are
- Ambitious and rigorous. You care more about being right than being clever.
- 3-6 years in ML or LLM engineering, with something in production.
- Strong on evaluation, not just prompting. Python.
- You read the papers and ship anyway.
What you get
- A founding role with real ownership and founding equity.
- Production data from global manufacturers, and engineers who tell you exactly where the model is wrong.
- A problem that matters: AI that engineers trust with safety-critical decisions.