Applied Healthcare Researcher
Protege · posted 1 hours ago
At a glance
- Location
- Remote
- Posted
- Aug 4, 2026
- Auto-expires by
- Aug 12, 2026 (if not re-listed at source)
Remote Score for Protege
Protege hasn't been scored yet. Read the rubric at how we score.
About this role (from Protege's posting)
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
Role Overview
We are hiring Applied Healthcare Researchers to join a team within DataLab focused entirely on healthcare training data.
Our customers are researchers at the frontier labs and AI startups building specialized healthcare models. They come to us with model-development problems, not dataset specifications. Figuring out which healthcare data actually solves their problem, and proving that it does, is the research question we answer in DataLab.
In this role you will work directly with researchers at those labs to understand what they're trying to train or evaluate, determine what healthcare data can support it, and do the research needed to demonstrate that it will. This is fast-iterating, customer-facing research on a customer's timeline. You will be the primary technical and research link to the customer — not a technical resource brought in for credibility, but the person driving the conversation and pulling in the solutions, engineering, and data partnerships teams as needed.
Core Responsibilities
Customer Research Partnership
You will be the research partner to AI researchers at frontier labs and startups who are working on healthcare problems.
• Serve as the primary technical and research point of contact for healthcare customer conversations.
• Translate a lab's model-development goals into concrete, feasible data strategies.
• Help customers scope opportunities and identify the highest-value data available to them.
• Explain data limitations, tradeoffs, and potential biases to technically sophisticated stakeholders while grounding conversations in what real-world data actually looks like.
• After delivery, answer the research questions customers raise about the data we provided. Delivery is not the end of the relationship.
Applied Research & Method Development
Curating the right data product is a research problem, and you'll own solving it.
• Develop and evaluate methods — fine-tuning, LLM-based extraction, classification, rules-based approaches, or whatever the problem calls for — to demonstrate that a dataset can support a customer's training or evaluation objective.
• Design and run feasibility research pre-contract: can this data support this model objective, at what quality, with what caveats.
• Build the evidence base that makes a data strategy credible — benchmarks, validation analyses, error characterization, and honest assessments of where the data falls short.
• Partner with the Assessments team on healthcare benchmarks across modalities.
Data Feasibility & Dataset Strategy
• Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data.
• Identify proxy variables or alternative dataset structures when the ideal variable doesn't exist.
• Analyze partner and source datasets — schema, field availability, quality, completeness, and required transformations.
• Contribute to our point of view on which healthcare data matters most for which modality and which stage of model development.
• Help evaluate new data partners and identify datasets worth acquiring before a customer asks for them.
Reusable Research & Scaling
• Produce reusable research, evidence, and technical collateral rather than starting from scratch for each opportunity.
• Identify where a successful one-off approach should become a repeatable workflow, and work with Product and Engineering to operationalize it.
• Help expand proven healthcare datasets across multiple customers instead of selling them once.
Cross-Functional Collaboration
• Work with Solutions and FDEs from the beginning of an opportunity.
• Coordinate with Healthcare Data Partnerships on sourcing and with Product and Engineering on tooling.
Required Experience & Skills
• Advanced degree (PhD or Master's plus 3+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience.
• Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data.
• Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training.
• Strong Python and SQL, with the ability to work independently against large datasets.
• Experience designing evaluations — measuring data quality and dataset representativeness.
• Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans.
• Comfort operating on a customer's timeline without lowering the standard of the research.
Ideal Profile
The ideal candidate:
• Is energized by working directly with customers, and specifically by working with other researchers as peers.
• Moves fast on messy, real-world problems and knows which corners can and cannot be cut.
• Is rigorous about what the data can and cannot support, and willing to tell a customer when the answer is no.
• Enjoys the full arc — scoping a vague problem, doing the research, and showing the result to the person who asked for it.
• Thinks about leverage: builds the reusable version rather than the one-off when it's worth doing.
About DataLab
DataLab exists because truly useful data is rare — and the frontier of AI development only moves forward when high-quality data makes it possible.
We believe data is one of the most underdeveloped layers of the AI stack. Our work focuses on building and evaluating high-value datasets grounded in real-world workflows and economically meaningful tasks. Our research spans data quality, evaluation design, privacy-preserving transformation, and task-grounded AI training data.
Protege Values
Pass the Loved Ones’ Test
We act with integrity and do the right thing — especially when it’s hard and no one is watching.
Always Find a Way
We are resourceful, resilient builders who solve hard problems and push through obstacles.
Go Fast and Grow Fast
Velocity matters. We move with urgency, learn quickly, and continuously improve as individuals and as a company.
Practice Kindness and Candor
We communicate directly and respectfully, building trust through honest feedback and genuine care for one another.
Deliver Together
We win as one team. Collaboration, accountability, and shared ownership drive our success.
Own the Outcome. Hone the Craft.
We take pride in our work, sweat the details, and continuously raise the bar for excellence.