H Company Logo

H Company

Senior Software Engineer - London

Posted 5 Hours Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Senior level
Hybrid
London, Greater London, England, GBR
Senior level
Build and operate a scalable evaluation framework for computer-use agents and AI models. Responsibilities include benchmark integration, orchestration, scheduling, observability, reproducible trials, distributed systems operation, API and external-service integrations, customer collaboration, and framework standards. The role requires production backend development in Python, Kubernetes and public-cloud experience, database and messaging knowledge, and end-to-end system ownership.
The summary above was generated by AI
About H:

When we released Holo4 on 28 September, we published every trajectory behind its public benchmark scores at trajectories.hcompany.ai. For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these.

H builds computer-use agents and the models behind them. Developers use them through a managed API, and our forward deployed engineers take them into enterprise workflows.

What this team owns

The evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up.

What you'd be doing
  • Integration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time.

  • Setting the standard for how a benchmark enters the framework, and building the checks that enforce it.

  • Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue.

  • Reproducible results across trials, so a release decision rests on numbers that hold.

  • Whatever stack a benchmark calls for. One week that's cluster tuning; the next it's a browser extension or desktop environments.

  • Time with customers, from single developers to large companies, to find out what they want measured, then automating it so the results flow back into our harnesses and models.

The first few months

By 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade-offs of the whole evaluation system, and be the person the rest of H asks about evaluations.

Who you'd work with

Ceiran Chapman, our VP Engineering, is hiring for this role. You'd join the evaluation team. The people relying on your work day to day are H's researchers and forward deployed engineers.

What we think it takes

Likely a good fit if you

  • Have spent 5+ years in backend development, with production Python at the core, and use coding agents to go faster without letting quality drop.

  • Have built test, QA or evaluation tooling that other teams depended on, and care whether a number is right.

  • Have operated distributed systems on Kubernetes in a public cloud. AWS experience helps most.

  • Have built and shipped systems end to end, including APIs (REST or GraphQL) and integrations with outside services.

  • Know relational and non-relational databases, and message queues such as SQS, RabbitMQ or Kafka.

  • Instrument what you build, with metrics, tracing and monitoring from the start.

Stronger still if you have

  • Measured LLM quality before, or built agents yourself.

  • Packaged and run workloads in Docker and on virtual machines.

  • Used Temporal, Dask, FastAPI, PostgreSQL, Grafana or Datadog.

  • Automated web or desktop software with Playwright, Selenium or a browser extension you wrote.

  • Set standards other engineers follow, through code review, design review or mentoring.

You do not need a background in machine learning. We'll work that out with you. If you match most of this but not all of it, apply anyway.

How we hire

A 30 minute call with our Talent team, a 60 minute technical challenge, a 60 minute system design interview, and a 30 minute final conversation with Ceiran. About 3.5 hours in total.

Practicalities

London posting: Hybrid in London. That means 3 office days a week and a Paris trip about once every 4 to 6 weeks. There is a Paris posting for the same role. We offer a competitive package.

Similar Jobs

3 Hours Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
eCommerce • Information Technology • Marketing Tech • Software
Design, build, and operate Akeneo’s Supplier Data Manager at enterprise scale. Develop reliable data-processing pipelines for messy supplier catalogues, including deterministic and LLM-powered extraction and classification. Create collaboration and governance workflows, ensure production observability and operational excellence, participate in customer discovery, and collaborate across teams. The role requires strong software engineering fundamentals, architectural thinking, debugging ability, end-to-end ownership, and clear communication.
Top Skills: Claude CodeGCPInfrastructure As CodeLlmsPHPPythonReactTypescript
3 Hours Ago
Hybrid
London, Greater London, England, GBR
Entry level
Entry level
eCommerce • Information Technology • Marketing Tech • Software
Build and ship end-to-end features for Akeneo Supplier Data Manager across data-processing pipelines, backend services, and TypeScript/React frontend systems. Help process large, messy supplier catalogues using deterministic workflows and LLM-based extraction and classification. Own features through production monitoring and support, participate in customer discovery, collaborate across the Product Cloud, and develop engineering skills through pairing, code review, and mentorship.
Top Skills: Claude CodeJavaScriptLarge Language ModelsPHPPythonReactTypescript
4 Hours Ago
Remote or Hybrid
Entry level
Entry level
Cloud • Healthtech • Social Impact • Software • Biotech
Drive new business and revenue growth across the Nordic territory by generating pipeline, managing complex SaaS sales cycles, forecasting accurately, presenting solutions, negotiating with technical and executive stakeholders, and closing strategic accounts. The role also involves account relationship management, collaboration with marketing, product, customer success, and channel teams, Salesforce data integrity, and use of structured sales methodologies such as MEDDICC.
Top Skills: Artificial IntelligenceMeddiccSaaSSalesforce

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account