Unlock Full Resume Report

New offer - be the first one to apply!

September 14, 2026

Data Scientist II - Big Data R&D, Identity Graph & Deceased Monitoring

Mid

140,000 - 170,000 USD/yr

Seattle, WA

Quick Facts

  • Big Data R&D role focused on graph-based identity graph and entity-resolution capabilities
  • Builds algorithms and data pipelines for identity verification, deceased monitoring, and compliance

Description

Join the Big Data R&D team to develop graph-based algorithms and data pipelines on massive PII datasets. You will support modelers with high-quality features, evaluate new data sources, and work with senior data scientists and engineers on large-scale ML, distributed systems, and graph analytics.

Responsibilities

  • Design and implement machine learning, data mining, statistical, and graph-based algorithms for very large datasets
  • Analyze large datasets to develop and refine entity-resolution and identity-matching algorithms
  • Build and maintain data-processing pipelines (ETL, feature generation, normalization) with Spark/PySpark and AWS (EMR, S3)
  • Support senior data scientists with feature engineering, data exploration, error analysis, and A/B test setup
  • Evaluate new third-party and internal data sources via profile data quality checks and offline experiments
  • Implement and maintain SQL and Python/R code for data extraction, transformation, and validation; contribute to code reviews and basic testing
  • Provide analytical support to compliance/regulatory product teams via investigations, dashboards, and data deep dives
  • Communicate findings clearly and structuredly to peers and cross-functional partners
  • Work effectively in a fast-paced, cross-functional environment with ownership and follow-through

Requirements

  • Master’s degree with 2+ years experience, or Ph.D. with 1+ years experience, or equivalent practical experience
  • Located within 45 miles of a talent hub
  • No sponsorship available
  • Proficiency in Python or Scala
  • Solid experience writing and optimizing SQL for large datasets; comfort with data lake/warehouse environments
  • Hands-on experience with Spark or PySpark and common ML libraries (scikit-learn, XGBoost; TensorFlow/PyTorch a plus)
  • Familiarity with supervised/unsupervised ML and basic statistics (similarity measures, clustering, evaluation metrics)
  • Familiarity with UNIX environments and AWS ecosystem (EMR, S3); Databricks a plus
  • Exposure to graph techniques or graph databases (Neo4j, AWS Neptune, GraphFrames) is a strong plus
  • Bonus: experience with Elasticsearch or DynamoDB; workflow tools such as Airflow for automating data pipelines
  • Ability to break down loosely defined problems, ask clarifying questions, and iterate quickly with feedback

Benefits

  • Compensation range: $140K - $170K
  • Equal opportunity employer

Similar jobs you might like