Founding Applied Scientist
Build the reasoning and data-valuation systems that turn months of literature review into seconds, across loosely connected fields of research.
About this role
Sphere's focus is novel reasoning over proprietary research. We help scientists and engineers compress months of literature review into seconds, pulling from disparate academic fields that may be only loosely related. At the core of this is understanding what data is valuable: which data we retrieve, what data we make available to retrieve, and, over time, a value score that can be used to price data itself. You will own the systems that make that reasoning accurate, grounded, and trustworthy, and help shape how we measure and value the research flowing through Sphere.
What you'll own
- Reasoning quality over proprietary research, from retrieval through grounded, cited synthesis
- Post-training methods that specialize models for research reasoning: fine-tuning, LoRA and other PEFT, preference optimization (DPO / RLHF), and distillation
- Retrieval, ranking, and reranking that surface relevant work across loosely related fields
- Chunking, parsing, and representation quality for complex scientific and technical documents
- Benchmarking and evaluation frameworks that reflect how researchers actually work
- Data valuation: which data we retrieve, what we make available to retrieve, and a value score that can price data
- Citation fidelity, groundedness, and hallucination reduction
What you'll do
- Apply and combine post-training methods (fine-tuning, LoRA / PEFT, RAG, DPO / RLHF, distillation) to make reasoning over research sharper and better grounded
- Improve retrieval and cross-domain synthesis so a query surfaces the right work even from loosely related fields
- Design and run evaluation frameworks grounded in real research workflows, and define what "better" means in measurable terms
- Build the data-valuation layer: score which sources are worth retrieving and worth acquiring, growing toward a price signal for data
- Test and refine chunking, parsing, embedding, reranking, and selection strategies over scientific corpora
- Ship production improvements alongside product and infrastructure, not just experiments
What we're looking for
- Strong applied data science or applied ML background across information retrieval, ranking, NLP, or reasoning systems
- Hands-on with post-training methods: fine-tuning, LoRA and other PEFT, RAG, preference optimization (DPO / RLHF), or distillation
- Ability to design evaluation frameworks, not just run existing benchmarks
- Comfort turning fuzzy questions like "what data is valuable?" into measurable models and metrics
- Strong engineering instincts and willingness to ship production code
- Comfort with messy scientific and technical documents and real-world constraints
Strong pluses
- Scientific, technical, or multimodal research corpora
- Retrieval-augmented reasoning, agentic retrieval, or grounded-answer systems
- Citation-grounded evaluation, provenance, or source attribution
- Data valuation, pricing models, or value-of-information and active-learning approaches
- Optimizing training or inference under constrained hardware
Your first 90 days
- Establish a clear evaluation baseline for research reasoning, retrieval, and citation quality
- Ship at least one post-training or retrieval improvement that moves a measurable quality metric
- Stand up the first version of a data-valuation signal: which sources are worth retrieving and acquiring
- Lay out a practical roadmap toward a value score that can inform how we price data
Compensation
$165K - $220K base salary, plus 0.8% - 1.5% equity and benefits.
Individual pay is determined by experience, relevant education, and/or training.
Benefits
- Health, dental, and vision insurance
- Flexible paid time off
- Company-provided equipment
As an early-stage company, we're committed to building a benefits package that grows with us.
To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.
Sphere is an Equal Opportunity Employer; employment with Sphere is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.