Expert proficiency in Python and SQL for data engineering tasks
Strong experience with modern data engineering frameworks, including Apache Spark, Kafka, and Airflow
Deep expertise in extracting and processing data from complex ERP environments, specifically SAP S/4HANA and SAP Ariba, with familiarity in SAP BTP as a plus
Comprehensive understanding of Data Lakehouse architectures, such as Databricks and Delta Lake, for scalable and secure data storage
Experience with relational databases, including PostgreSQL, and vector databases like Weaviate and Milvus for AI-driven applications
Proven ability to develop data pipelines for Retrieval-Augmented Generation (RAG) solutions, conversational agents, and classical machine learning models using tools like dbt, Dagster, or Prefect
Proficiency in containerization technologies, including Docker and Kubernetes, and experience implementing CI/CD pipelines for secure data workflow deployment
Knowledge of defense-grade security protocols, including Role-Based Access Control (RBAC), audit logging, and data redaction policies for compliance with export controls and on-premise security requirements
Experience building and optimizing pipelines for high-performance GPU clusters and air-gapped environments
Familiarity with automated data quality frameworks to validate critical datasets such as Bill of Materials (BOM) and cost data for AI model accuracy
Ability to design and implement standardized procurement data models and taxonomies to harmonize fragmented datasets across multiple entities
Experience engineering pipelines for unstructured data ingestion, including PDF tender documents, CAD metadata, and historical CONOPS, transforming them into vector-ready formats for AI applications
Knowledge of Model-Based Systems Engineering (MBSE) principles and the ability to establish data lineage and traceability protocols for the Digital Thread
Responsibilities
5+ years of experience in Data Engineering, with at least 2 years focused on building pipelines for Machine Learning or Generative AI applications in an enterprise setting.
Expert proficiency in Python, SQL, and modern data engineering frameworks including Apache Spark, Kafka, and Airflow.
Strong experience extracting data from complex ERP environments, specifically SAP S/4HANA and SAP Ariba, with familiarity in SAP BTP as a plus.
Deep understanding of Data Lakehouse architectures (Databricks/Delta Lake), Relational Databases (PostgreSQL), and Vector Databases (Weaviate/Milvus).
Experience building pipelines for RAG solutions, Conversational AI agents, and classical ML models using tools such as dbt, dagster, or prefect.
Proficiency with containerization (Docker, Kubernetes) and CI/CD pipelines for deploying data workflows in secure environments.
Experience in Supply Chain, Manufacturing, or Defense sectors, with the ability to understand 'Bill of Materials' (BOM) structures and procurement lifecycles.
Ability to navigate the governance challenges between agile data work and rigid systems engineering requirements, ensuring data deliverables meet formal Stage Gate reviews.
Proven ability to collaborate with Data Scientists and Backend Engineers to define data schemas that support predictive modeling and AI agents.