Solve Intelligence Logo

Solve Intelligence

Data Engineer

Posted 5 Days Ago
Be an Early Applicant
In-Office
London, Greater London, England, GBR
Entry level
In-Office
London, Greater London, England, GBR
Entry level
Build and operate large-scale data ingestion, document processing, search, and serving systems for AI products. Responsibilities include handling massive heterogeneous datasets, designing schemas and indexes, enabling keyword, vector, and structured search, linking information across sources, and optimizing performance, reliability, latency, and cost. The role owns systems end to end, from source acquisition through user-facing functionality, and collaborates with AI researchers and product engineers.
The summary above was generated by AI
💡 About Us

We're the fastest-growing startup transforming the IP industry.

  • Traction: 20-30% MoM revenue growth; selling to 700+ global IP teams (DLA Piper, tech boutiques, and global enterprises).

  • Proven Value: Users report 50-90% efficiency gains using our AI platform.

  • Backing: Recently featured in Sifted following our $40M Series B announcement, bringing our total funding to $55M from elite investors including Y Combinator, 20VC, Visionaries, Microsoft and Thomson Reuters.

🏗️ About the role

We’re hiring a data engineer to build the ingestion and search systems behind Solve Intelligence’s AI products.

Our sources include global patent literature, case law, technical standards and contributions, scientific databases, academic papers, and content from across the web. The data spans structured records, documents, images, audio and video. You’ll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our products

You’ll own systems from source acquisition through to serving queries. The work includes:

  • Large-scale ingestion. Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.

  • Document processing and data quality. Extract useful content from complex documents and other formats. Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.

  • Search and serving. Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.

  • Connecting information across sources. Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.

  • Performance engineering. Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale. Diagnose CPU, memory and storage I/O bottlenecks, and tune jobs and infrastructure for throughput, latency and cost.

You’ll work closely with our AI researchers and product engineers, with substantial freedom to choose the approach and build the systems yourself.

🛠️ What you bring

Must haves:

  • Strong Python and SQL, with experience designing and operating production databases.

  • Solid experience building and operating production data pipelines over large, messy datasets.

  • Expertise with running search systems over large document collections.

  • End-to-end ownership from raw data to user-facing functionality.

  • A good understanding of schema design, indexing and query optimisation.

  • A track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement.

Nice to Have:

  • Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++ is useful.

👋 The Founders

You'll partner with a founding team of AI PhDs and elite systems engineers:

  • Sanj (CRO): PhD in AI (Gatsby Unit, UCL), ex-Huawei R&D, former lead at Magic Carpet AI (acquired).

  • Chris (CEO): PhD in AI (UCL), published researcher, ex-Dyson and Alan Turing Institute.

  • Angus (CTO): MEng Computer Science, ex-Qualcomm and Coremont (Brevan Howard).

What we offer
  • Competitive Salary + Significant Equity: We want you to have true ownership in the success of the company.

  • Founding Impact: You'll have a direct hand in how we build out the data infrastructure the rest of the product depends on.

  • Support: Full visa sponsorship and private medical insurance.

  • The Environment: Free meals and a seat at the table with an incredibly smart, ambitious team.

 

Similar Jobs

4 Days Ago
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Designs, builds, and operates Python-based data platforms, web applications, APIs, and agentic workflows. Develops batch and streaming pipelines, AWS data services, secure backend systems, CI/CD, monitoring, and governance controls. Translates business needs into production solutions, establishes reliable AI-agent safeguards, mentors engineers, reviews designs and code, and coaches business users building front ends with AI-assisted tools.
Top Skills: Amazon AthenaAmazon Aurora PostgresqlAmazon RedshiftAmazon S3AWSAws GlueAws LambdaCi/CdClickhouseDbtFastapiFlaskInfrastructure As CodeJavaScriptPostgresPrefectPythonStreamlitTypescript
5 Days Ago
Hybrid
Entry level
Entry level
eCommerce • Food • Logistics • Retail • Sales
Lead internal data engineering and tracking for a six-month contract. Responsibilities include overseeing an external engineering agency, reviewing architecture and dbt models, auditing end-to-end data flows, implementing automated data quality tests, monitoring website tracking, optimizing BigQuery, dbt, and Fivetran costs, and ensuring GDPR compliance. The role requires advanced SQL, BigQuery, dbt, GTM, GA4, Python, and Git expertise, plus the ability to lead technical delivery and support data stakeholders.
Top Skills: BigQueryDbtFivetranGdprGitGitGoogle Analytics 4 (Ga4)Google Cloud Platform (Gcp)Google Tag Manager (Gtm)PythonSQL
12 Days Ago
In-Office
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Develop and maintain regulatory reporting systems across SQL Server and GCP BigQuery. Build automated ETL/ELT pipelines, transform data into regulator-ready outputs, and support migration to cloud-native solutions. Implement data modelling, validation, monitoring, alerting, QA, documentation, and performance improvements while participating in code reviews and applying infrastructure as code, CI/CD, orchestration, and AI-assisted development practices.
Top Skills: Ci/CdClaude CodeCloud ComposerCloud FunctionsCloud StorageEltETLGitGitlabGoogle BigqueryGoogle Cloud Platform (Gcp)Pub/SubSQLSQL ServerTerraform

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account