Senior Clinical AI Data Lead
About kaiko
Clinicians and nurses work under heavier pressure than ever. More patients, more data, more decisions to make. kaiko give them back the focus, time, and headspace they need to give patients the best possible care.
kaiko is a European clinical AI lab, and one of the only companies working across the full stack. Our approach spans three layers: our own frontier models, a clinical AI workspace built for complex hospital needs and workflows ranging from preparation and decision support to diagnostics and interpretation.
Our product is in daily use at leading European hospitals, reducing preparation time and cognitive workload for clinical teams. The next decade of European healthcare will be shaped by AI. We're here to shape it in the best way possible. kaiko is a well-funded company with a growing international team, operating from Amsterdam and Zurich.
About the role
At Kaiko, we build AI that healthcare professionals (HCPs) use in their daily work. Clinical judgement is central to its quality. HCPs review, rate and challenge what our AI produces, helping ensure that what reaches the clinic is safe, reliable and genuinely useful.
You own the clinical expertise behind that work. You build the right network of HCPs, match the right clinicians to the right tasks, and turn their judgement into high-quality evaluation and annotation datasets that our AI team can rely on.
You sit inside the Clinical AI Engineering team, working closely with ML engineers and clinical experts. You help shape the tasks and share responsibility for clinical data quality.
A typical request might be:
400 oncologist-rated evaluations of a model behaviour, against a defensible rubric, delivered in three weeks.
Several hundred Dutch to English clinical translations, each verified by a bilingual HCP.
The role is based in Amsterdam and works in Dutch and English, with around half your time in the office. The network starts in the Netherlands and will grow from there.
What you'll own
The expert network. Find, screen and onboard the right HCPs, from students to specialists. Build relationships across Dutch UMCs, oncology centres and beyond, and know which level of expertise each task actually requires.
Programs. Turn requests from Clinical AI, Research and Product into well-run evaluation and annotation programs. Define the people, process, scale, timeline and cost needed to deliver them well.
Quality. Help design rubrics and guidelines, calibrate annotators, measure agreement, run adjudication and spot bias before it becomes part of the dataset.
Representativeness. Work with ML engineers to make sure evaluation data reflects how the product is actually used, and evolves as that usage changes.
Operations. Scale the pool up and down across levels, with a stable core on zero-hour contracts plus contractors as needed. Manage the practical side of staffing, access, contracts, budgets and timelines together with Recruiting, HR and Legal.
About you
Clinical fluency. You understand healthcare well enough to know who should do a task, how difficult it is, and whether the output is credible. You may come from a hospital, general practice, medtech, clinical research, CROs, trials, medical affairs or another healthcare environment.
Dutch and European healthcare. You are fluent in Dutch, know the Dutch healthcare system well enough to build relationships with clinicians and institutions, and are comfortable extending that network across Europe and beyond.
Program and operations leadership. You have led complex, multi-contributor programs to a deadline, ideally involving clinical data, research, annotation or evaluation.
Annotation and evaluation methodology. You are comfortable with rubric design, calibration, sampling, inter-annotator agreement and adjudication. You understand how a poorly designed rating task can destroy signal.
Winning experts over. You find and win over the right experts fast, and know who to approach for what. Clinicians trust you because you speak their language and understand how they work.
Cross-functional communication. You can move easily between HCPs, ML engineers, product and operations, turning draft requests into something concrete and usable.
Nice to have
External vendors. Experience commissioning or managing annotation vendors.
International networks. Experience building expert networks across European markets or beyond.
LLM evaluation. Familiarity with concepts such as LLM-as-judge, reward-model data, and failure modes such as sycophancy, verbosity, and reward-hacking.
We are excited to gather a broad range of perspectives in our team, as we believe it helps us build better products for a broader set of people. If you're excited about us but don't fit every single qualification, we still encourage you to apply: we've had incredible team members join us who didn't check every box.
Why kaiko
At kaiko, we believe the best ideas come from collaboration, ownership and ambition. We've built a team of international experts where your work has direct impact. Here's what we value:
Ownership: You'll have the autonomy to set your own goals, make critical decisions, and see the direct impact of your work.
Collaboration: You'll approach disagreement with curiosity, build on common ground and create solutions together.
Ambition: You'll be surrounded by people who set high standards, see obstacles as opportunities, and work relentlessly to create better outcomes for patients.
In addition, we offer:
An attractive and competitive salary, a good pension plan and 25 vacation days per year.
Great offsites and team events to strengthen the team and celebrate successes together.
A EUR 1000 learning and development budget to help you grow.
Autonomy to do your work the way that works best for you, whether you have a kid or prefer early mornings.
An annual commuting subsidy.
Our interview process
Designed to assess mutual fit across skills, motivation, and values. It typically includes:
Screening call: motivation, career goals, and initial fit.
Working session (take-home or live): given a sample clinical task, design an end-to-end plan to produce trustworthy data on a deadline.
Review discussion: walk through your submission and the trade-offs you made.
Onsite: present a data, research or annotation program you ran end to end, followed by conversations with ML, clinical, and ops colleagues.
- Department
- Engineering
- Locations
- Amsterdam
- Remote status
- Hybrid