Develop, deploy, and maintain LLM pipelines and RAG QA/search systems; design and optimize prompts and multi-agent LLM architectures; operate multi‑GPU/cluster inference; build evaluation pipelines for model quality, bias, and hallucination; collaborate with product and CS teams to integrate conversational AI.
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.
We are seeking a highly skilled professional to join our team, focusing on advancing through innovative AI solutions.
The successful candidate will develop and refine Large Language Models (LLMs) to extract actionable insights, improve business decision-making, and optimize prompt design for more accurate outputs. Additionally, the role includes creating scalable and robust LLM/RAG frameworks tailored to customer service scheduling, fostering innovation and maintaining a competitive market edge.
This role is 100% Remote, Work from Home based.
Responsibilities
- Own the full LLM pipeline from data preparation to production real case usage.
- Design, iterate and optimize prompts (zero-/few-shot, chain-of-thought, tool-calling, etc.) to maximize model utility and safety across products and languages.
- Build and maintain Retrieval-Augmented Generation (RAG) QA/search systems that connect to multi-source knowledge bases.
- Familiar with vLLM/SGLang inference architectures and have proven experience deploying and operating LLM services on multi‑GPU or cluster environments.
- Design, implement and operate multi‑agent LLM architectures (e.g. LangGraph, CrewAI, AutoGen) including task decomposition, agent orchestration, memory sharing and tool‑calling workflows.
- Develop evaluation pipelines (automatic metrics & human feedback) to measure prompt and model quality, bias, and hallucination rates.
- Collaborate with product and CS teams to integrate AI models into conversational Chatbot in different scenarios.
- Track cutting-edge research, author tech blogs, and keep improve current architecture.
Requirements
- Master’s Degree or higher in Computer Science, Data Science or related field..
- At least 2 years of deep-learning/NLP experience, including 1+ year practical LLM work (SFT, DPO, RAG, quantization, inference optimization, etc.).
- Demonstrated prompt engineering & tuning expertise (few-shot design, structured prompting, prefix-/p-tuning, reward re-ranking, safety filtering).
- Practical experience building and deploying multi‑agent LLM workflows, with understanding of agent‑orchestrator patterns, shared memory, long‑horizon planning and guard‑rail design.
- Proficient in both English and Chinese communication for efficient cross team collaboration
Why Binance
• Shape the future with the world’s leading blockchain ecosystem
• Collaborate with world-class talent in a user-centric global organization with a flat structure
• Tackle unique, fast-paced projects with autonomy in an innovative environment
• Thrive in a results-driven workplace with opportunities for career growth and continuous learning
• Competitive salary and company benefits
• Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)
Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.
By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice.
Similar Jobs
Angel or VC Firm • Artificial Intelligence • Software
Lead design and build of an end-to-end AI contract-review agent: prompt orchestration, RAG pipelines, fine-tuning, web UI and tooling, data/model infra, deployment (CI/CD/cloud), and product-driven features like assumption mapping and contract diagnostics.
Top Skills:
AutogenAWSCi/CdClaudeFine-TuningGCPLangchainLlmsNext.JsOpenaiPineconePostgresPrompt EngineeringRagReactSupabaseVector StoresVercelWeaviate
Angel or VC Firm • Artificial Intelligence • Software
Lead product vision and discovery for a genomics-driven diagnostic and personalization platform. Validate hypotheses with patients, clinicians, and partners; translate research into product specs; coordinate labs and data providers; partner with engineers and researchers to build AI prototypes; shape data, regulatory, and GTM strategy while making core product and architectural decisions.
Top Skills:
AIClinical Trial DataData PrivacyEhrGenomic DataGenomicsGxpHipaaIrbMachine LearningPhenotypic Data
Angel or VC Firm • Artificial Intelligence • Software
Build core components of a diagnostic and personalization platform: backend data pipelines, frontends, services for genomic data ingestion, ML model integration, APIs, and cloud infrastructure. Collaborate with product and ML teams, iterate on prototypes and MVPs, and make architectural decisions supporting scalability and reliability.
Top Skills:
AIAPIsAWSCloud InfrastructureData PipelinesFastaGCPMachine LearningVcf
What you need to know about the London Tech Scene
London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

