Unlock Full Resume Report

New offer - be the first one to apply!

September 16, 2026

Senior Data Scientist, Biologics Discovery

Senior • On-site

109,000 - 174,800 USD/yr

Titusville, NJ

Quick Facts

  • Role: Senior Data Scientist (Biologics Discovery)

  • Team: Data, Data Science & Artificial Intelligence (DDSAI) partnering with In Silico Discovery (ISD)

  • Work setup: On-site at Spring House, PA (strongly preferred), Titusville, NJ, Raritan, NJ, or Madrid, Spain (no remote)

Description

You will build data-facing machine learning capabilities that make biologics data model-ready and trustworthy for molecular-property modeling. The role focuses on featurization, model-ready dataset curation, evaluation frameworks, and applied models on assay and sequence data. You will collaborate across data/infrastructure partners, ISD, and discovery scientists to strengthen earlier prioritization, risk flagging, and hypothesis generation.

Responsibilities

  • Develop featurization pipelines and model-ready datasets from antibody/protein sequence, construct, assay, and biophysical data

  • Define feature/label/aggregation levels with data engineers while preserving raw representations where needed

  • Curate, document, and version datasets for reproducibility and traceability

  • Build and evaluate applied ML models on biologics assay, biophysical, and sequence/construct data

  • Define evaluation frameworks that reduce model risks and support trustworthy performance

  • Hand off standardized, traceable training datasets with ISD and align on ownership boundaries between data science and modeling

  • Partner with discovery scientists to frame ML problems around real decision points in the DMTL cycle

  • Work with ontology and MLOps colleagues to ensure consistent semantics and reliable model movement into use

  • Champion reproducibility, documentation, and responsible AI

Requirements

  • Master’s or Ph.D. in a relevant field (Computer Science, Machine Learning, Computational Biology, Bioinformatics, Statistics, or related)

  • 2+ years of applied ML experience including model development, evaluation, and dataset curation on complex scientific/biomedical data

  • Strong Python, modern ML stack (PyTorch, scikit-learn), and SQL

  • Experience transforming complex heterogeneous experimental data into robust features and training sets, including exposure to cloud training and data infrastructure

  • Knowledge of evaluation/validation and risks of leakage and distribution shift

  • Ability to collaborate effectively with experimental scientists and modeling partners

Benefits

  • Performance-based annual bonus eligibility

  • Medical, dental, vision, life insurance, short- and long-term disability, business accident insurance, and group legal insurance

  • Retirement plan and savings plan (401(k))

  • Time off: up to 120 hours vacation, up to 40 hours sick time, and holiday pay including floating holidays

Similar jobs you might like