Senior Software Engineer - Paris

H Company · Paris, Île-de-France

ABOUT H:

When we released Holo4 https://hcompany.ai/newsroom/holo4 on 28 September, we published every trajectory behind its public benchmark scores at trajectories.hcompany.ai http://trajectories.hcompany.ai. For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these.

H builds computer-use agents and the models behind them. Developers use them through a managed API https://hcompany.ai/newsroom/computer-use-agents-api, and our forward deployed engineers take them into enterprise workflows.

WHAT THIS TEAM OWNS

The evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up.

WHAT YOU'D BE DOING

THE FIRST FEW MONTHS

By 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade-offs of the whole evaluation system, and be the person the rest of H asks about evaluations.

WHO YOU'D WORK WITH

Ceiran Chapman, our VP Engineering, is hiring for this role. You'd join the evaluation team. The people relying on your work day to day are H's researchers and forward deployed engineers.

WHAT WE THINK IT TAKES

Likely a good fit if you

Stronger still if you have

You do not need a background in machine learning. We'll work that out with you. If you match most of this but not all of it, apply anyway.

HOW WE HIRE

A 30 minute call with our Talent team, a 60 minute technical challenge, a 60 minute system design interview, and a 30 minute final conversation with Ceiran. About 3.5 hours in total.

PRACTICALITIES

Paris posting: Hybrid in Paris. That means 3 office days a week and a London trip about once every 4 to 6 weeks. There is a London posting for the same role. We offer a competitive package.

Apply on H Company's site