Unlock Full Resume Report

New offer - be the first one to apply!

September 1, 2026

Principal AI Data Readiness Architect

Senior • On-site

18,000 - 25,000 PLN/yr

Krakow, MA, Poland

We are seeking a Staff/Principal AI Data Architect to modernize an enterprise data ecosystem for new AI and ML tools, including automated classification and summarization, agentic workflows, and RAG. This role focuses on data readiness, governance, quality, and secure access.

You will define the standards, contracts, and observability that make structured and unstructured data trustworthy, discoverable, and easy to consume in batch and near-real-time contexts. Orchestration tooling is still being determined, with a current focus on an Airflow-centric approach. The person in this role will help make decisions about data infrastructure implementation and tooling.

Strategy and Standards

  • Define the enterprise AI data architecture vision, principles, and reference architectures.
  • Lead cross-functional reviews with IT, security, legal/privacy, and business stakeholders to align on data readiness roadmaps.

Data Contracts, Catalog, and Modeling

  • Establish data contracts for AI consumption, including schemas, semantics, classifications, and SLAs, and govern schema evolution for backward compatibility.
  • Make the data catalog the system of record for lineage, ownership, definitions, and policy labels; integrate with intake and change management.
  • Define standard data models and semantic conventions that improve joinability and reuse across domains.

Data Quality and AI Data Observability

  • Implement an enterprise data quality framework and automated scorecards for freshness, completeness, accuracy, and consistency.
  • Monitor anomalies and schema drift; publish AI data readiness dashboards for catalog coverage, lineage depth, PII detection coverage, and contract adherence.

Pipelines, Orchestration, and Access

  • Standardize patterns for ingestion, processing, storage, serving, and environment promotion using Airflow or other standard ETL/orchestration tools and CI/CD for data workflows.
  • Define secure, consistent access patterns and APIs for downstream analytics and AI consumers.

Vector Search and RAG Readiness (Enablement)

  • Drive the foundational architecture and standards needed to enable Retrieval Augmented Generation (RAG) and semantic search capabilities across the enterprise.
  • Provide guidance for chunking and segmentation policies, deduplication, and hybrid search compatibility; downstream teams implement embeddings and vector stores.

Security, Privacy, and Compliance

  • Define safe-access patterns for AI consumption to prevent sensitive data exposure.
  • Enforce security baselines, including encryption, RBAC/ABAC, masking/tokenization, and policy-as-code for access.

Financial Operations

  • Architect transparent cost attribution and controls, including tagging, storage tiering, and retention, to enable informed cost and performance choices.
  • Assist leadership with recommendations to improve efficiency and automated triggers to identify planned budget allocation violations.

Collaboration and Mentorship

  • Provide reference templates for AI-ready datasets, contracts, and catalog usage; mentor engineers and analysts on best practices.
  • Collaborate with other system architects to ensure continued reliability and identify opportunities to improve the overall ecosystem.

Basic Requirements

  • 8+ years in data engineering, architecture, or platform roles, preferably with more than 1 year at the Staff/Principal level.
  • Expert SQL and Python, with a track record of building enterprise data governance, contracts, and quality frameworks.
  • Experience operating production data platforms in batch and near-real-time contexts with strong lineage and access control.
  • Practical unstructured data governance experience, including metadata standards, classification, and PII detection/redaction.
  • Hands-on experience with catalogs and lineage as systems of record for definitions, ownership, and policy.
  • Familiarity with vector/RAG readiness concepts, including schemas, metadata, and provenance, without owning embeddings or model development.
  • Experience with workflow orchestration, such as Airflow, and CI/CD/testing for data pipelines.

Similar jobs you might like