Founding ML Engineer, Evaluation & Agents

The hard problem in industrial AI isn’t generation. It’s judgment. Our agents draft and maintain FMEAs and score failure records for live enterprise customers, and every customer measures the system against their own best engineers.

We’re looking for a world-class ML engineer who wants to own that judgment: evaluation, calibration and the learning loop.

What you’ll do

  • Build evaluation sets with customer engineers and measure accuracy where it matters most, on the highest-risk and lowest-risk calls.
  • Calibrate scoring and reasoning across model providers, and swap models without losing accuracy.
  • Turn engineer corrections into memory the agents reuse.
  • Own quality, cost and latency for every deployment.

Who you are

  • Ambitious and rigorous. You care more about being right than being clever.
  • 3-6 years in ML or LLM engineering, with something in production.
  • Strong on evaluation, not just prompting. Python.
  • You read the papers and ship anyway.

What you get

  • A founding role with real ownership and founding equity.
  • Production data from global manufacturers, and engineers who tell you exactly where the model is wrong.
  • A problem that matters: AI that engineers trust with safety-critical decisions.
Job Category: Engineering
Job Type: Full Time
Job Location: Remote, Europe

Apply for this position

Allowed Type(s): .pdf, .doc, .docx
Cookie settings
Necessary cookies keep the site working and are always on. Analytics and advertising cookies are only set if you allow them. You can change this at any time. Privacy policy
Necessary
Allow Necessary cookies.
Analytics
Allow Analytics cookies.
Advertising
Allow Advertising cookies.