Canva Logo

Canva

PhD Research Scientist Intern - Edge AI

Posted 12 Days Ago
Be an Early Applicant
Hybrid
London, Greater London, England, GBR
Internship
Hybrid
London, Greater London, England, GBR
Internship
14-week PhD research internship to optimise, deploy, and evaluate video-capable vision-language models for on-device intelligent captioning. Responsibilities include model selection, quantisation/pruning/distillation, hardware-specific compilation, deployment on mobile hardware, comparative and human evaluation, documentation, and delivering a prototype plus patent and paper-ready results.
The summary above was generated by AI
Company Description

Hey, g'day, mabuhay, kia ora, 你好, hallo, vítejte!

Our global HQ is in Sydney, Australia, but our London campus sits in Hoxton Square, right in the middle of Shoreditch. It's a bit of a warren of stairs and rooms — you will get lost at first, and someone will happily give you a tour. It's a space where our UK team comes together to connect, create and collaborate.

Fun fact: our London team is one of the places where the AI powering Canva gets built.

This role is based in London, and we're looking for someone who calls it home. Our hybrid way of working gives you flexibility — you'll have the option to work from home as well as connecting and collaborating with your team in-person, on campus. We trust teams to choose the balance that empowers them to achieve their goals.

Job Description

At Canva, our mission is to empower the world to design. We're building AI that feels magical and lands real impact for millions of people, helping anyone create with confidence. We're looking for a research intern who is excited by efficient ML and edge deployment to help us bring video-capable vision-language models onto the devices in people's pockets.

About the team

We're the Video Storytelling team, working on the models and systems behind Canva's video AI experiences. We partner closely with our Edge AI group, who are building Canva's on-device inference capability, to explore what's possible when AI runs directly on users' own hardware. We already own several of the server-side capabilities this work builds on, so you'll be joining a team with a strong command of the data, models, and pipelines behind the problem.

About the role

This is a 14-week research internship focuses on one clear question: can a video-capable vision-language model be optimised to run efficiently on high-traffic consumer phones, while retaining enough capability to serve a real product use case?

The use case is intelligent captioning, where the model's visual understanding of a video drives context-aware, intelligently placed captions. You'll own the complete arc, from model selection through optimisation, deployment, and measurement, delivering a working on-device prototype plus a benchmarked map of what current consumer hardware can and can't do. You'll inherit mature data and pipelines from our work, so you can benchmark directly against a strong reference rather than building from scratch. You'll be supported by supervisors with deep on-device AI backgrounds, weekly 1:1s, and collaborators across teams.

What you'll do

  • Survey candidate video-capable VLMs (e.g. Gemma, Qwen-VL, SmolVLM, MiniCPM-V) and determine the best starting point

  • Apply model optimization techniques and architecture improvements to specialize vision-language models for on-device deployment, including quantization, pruning, distillation, hardware-specific compilation, and task-specific fine-tuning for caption placement.

  • Deploy the model on real, high-traffic mobile hardware through our on-device inference library, iterating the optimisation-deployment loop against real on-device measurements.

  • Run comparative evaluation against at least one alternative optimisation path, and human evaluation against our server-side captions quality bar.

  • Document your findings clearly enough that the team can act on them, mapping which workloads are viable on-device today and which aren't yet, and why.

  • Compile your output into a patent filing and a paper publication.

You're likely a match if you have

  • Strong Python and hands-on PyTorch experience, including training and fine-tuning vision-language models.

  • A solid understanding of modern vision-language and multimodal architectures, with the ability to pick up a recent paper and reproduce it.

  • Experience with optimisation methods like quantisation, pruning, or distillation, and a clear sense of what each costs you in accuracy.

  • Experience deploying models on-device or at the edge with runtimes like Core ML, LiteRT/TFLite, ONNX Runtime, or ExecuTorch, working within real memory and latency budgets.

  • Experience running your own research project end to end: making a plan, measuring carefully, and iterating on what you find.

  • Current enrolment in a PhD in ML, CS, or a related field, with first-author papers at venues like CVPR, NeurIPS, ICCV/ECCV, ICLR, or ICML.

Nice to have

  • Experience with video understanding models, ideally the token-efficient kind.

  • Publications or open-source contributions in efficient ML, multimodal models, or edge AI.

  • Experience writing custom kernels for inference optimisation.

  • Experience deploying models across different on-device hardware accelerators (e.g. Apple Neural Engine, DSPs).

  • Experience working across research and product teams, in industry or on a previous internship.

Additional Information

Other stuff to know

We make hiring decisions based on your experience, skills and passion, as well as how you can enhance Canva and our culture. When you apply, please tell us the pronouns you use and any reasonable adjustments you may need during the interview process.

We celebrate all types of skills and backgrounds at Canva so even if you don’t feel like your skills quite match what’s listed above - we still want to hear from you!

Please note that interviews are conducted virtually.

Canva London, England Office

London, United Kingdom

Similar Jobs

2 Hours Ago
Hybrid
London, Greater London, England, GBR
Senior level
Senior level
Artificial Intelligence • Fintech • Greentech • Sales • Software • Travel • Hospitality
Own Perk’s US and UK PR strategy, building Tier 1 relationships with business and financial media. Lead proactive pitching, crisis communications, corporate narrative development, executive profiling, media training, and translation of AI and technology developments into compelling stories. Manage PR agencies, measure coverage and ROI, and integrate AI tools into research, drafting, monitoring, and reporting. Partner closely with senior leadership, product, engineering, and global communications teams in a fast-paced, hybrid environment.
Top Skills: Artificial Intelligence (Ai)SaaS
2 Hours Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Build and maintain resilient systems, automation, operational APIs, observability, and infrastructure-as-code tooling. Diagnose distributed-system incidents from edge to origin, manage monitoring and alerting platforms, configure Cloudflare services, participate in incident response and post-mortems, and improve reliability, performance, and operational consistency. Collaborate across SRE, development, and IT Operations while mentoring colleagues and applying AI tools to increase productivity and root-cause analysis.
Top Skills: AnsibleCdnCloudflareCoding AssistantsDdos ProtectionDnsGoGrafanaInfrastructure As CodeJavaScriptLlm PlatformsNew RelicOpentelemetryPagerdutyPythonShell ScriptingSplunkTerraformWaf
2 Hours Ago
Hybrid
Entry level
Entry level
Digital Media • Gaming • Software • Esports • Automation
Leads the re-architecture and delivery of business-critical risk and regulatory platforms using Golang, React, cloud technologies, Kafka, SQL, .NET, and TypeScript. Provides technical leadership, governance, solution design, code-quality oversight, estimation, documentation, escalation support, and mentoring. Ensures highly available, scalable, maintainable, performant, and cross-browser-compatible systems while collaborating with stakeholders and architects from development through production.
Top Skills: .NetCloud PlatformsGoKafkaReactSQLTypescript

What you need to know about the London Tech Scene

London isn't just a hub for established businesses; it's also a nursery for innovation. Boasting one of the most recognized fintech ecosystems in Europe, attracting billions in investments each year, London's success has made it a go-to destination for startups looking to make their mark. Top U.K. companies like Hoptin, Moneybox and Marshmallow have already made the city their base — yet fintech is just the beginning. From healthtech to renewable energy to cybersecurity and beyond, the city's startups are breaking new ground across a range of industries.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account