Lead Data Engineer at OPENKYBER LLC
California, California, USA -
Full Time


Start Date

Immediate

Expiry Date

24 Dec, 26

Salary

120000.0

Posted On

07 Oct, 26

Experience

5 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Industry

Information Technology & Services

Description

Job Description:

Lead/Senior Databricks Data Engineer Position: Lead Databricks Data Engineer / Databricks Architect Experience: 10+ Years Location: Remote Job Summary We are looking for an experienced Lead Databricks Data Engineer / Databricks Architect with 10+ years of overall experience in Data Engineering and strong hands-on expertise in Databricks, Apache Spark, PySpark, SQL, Delta Lake, and cloud-based data platforms . The candidate will be responsible for designing and implementing scalable Lakehouse architectures , enterprise data pipelines, data integration solutions, governance frameworks, and high-performance analytics platforms using Databricks.

  • Key ResponsibilitiesDesign and develop scalable data engineering solutions using Databricks and Lakehouse architecture .
  • Build robust ETL/ELT pipelines using PySpark, Spark SQL, Python, and SQL .
  • Design and implement Bronze, Silver, and Gold/Medallion architecture .
  • Develop and optimize Delta Lake tables, including MERGE, schema evolution, Change Data Feed, and incremental processing.
  • Build batch and real-time/streaming pipelines using Structured Streaming, Auto Loader, and Lakeflow .
  • Develop and manage Databricks Jobs/Workflows for pipeline orchestration, scheduling, dependencies, retries, and monitoring.
  • Implement enterprise data governance using Unity Catalog , including access control, data lineage, auditing, catalogs, schemas, and external locations. Unity Catalog provides centralized governance, access control, lineage, and auditing across Databricks data and AI assets.
  • Perform Spark and Databricks performance tuning , including cluster configuration, partitioning, caching, query optimization, Photon, and workload optimization.
  • Design data models supporting Data Warehousing, BI, Analytics, and AI/ML workloads .
  • Integrate Databricks with cloud platforms such as AWS, Azure, or Google Cloud Platform .
  • Work with cloud services such as AWS S3, Azure ADLS Gen2, Azure Data Factory, AWS Glue, Synapse, Event Hubs/Kafka/Kinesis , as applicable.
  • Implement CI/CD pipelines using Git, Azure DevOps/GitHub/Jenkins and Databricks deployment capabilities.
  • Work with Terraform/IaC for infrastructure provisioning and automation.
  • Troubleshoot production pipeline failures, performance issues, data-quality problems, and Spark/cluster issues.
  • Establish data quality, monitoring, logging, and observability practices.
  • Provide technical leadership, code reviews, architecture guidance, and mentorship to junior/mid-level engineers.
  • Collaborate with Data Architects, Data Scientists, Business Analysts, DevOps teams, and application teams.

Required Technical SkillsDatabricks, Databricks Lakehouse Platform, Delta Lake, Unity Catalog, Databricks Workflows/Jobs, Lakeflow / Delta Live Tables, Auto Loader, Databricks SQL, Databricks notebooks, Databricks Asset Bundles, Photon, Cluster/workload optimization, Big Data, Apache Spark, PySpark, Spark SQL, Structured Streaming, Kafka, Batch and real-time data processing Programming Python, SQL, PySpark, Scala good to have Cloud Strong experience in at least one AWS: S3, Glue, EMR, Lambda, Redshift, IAM, Kinesis Azure: ADLS Gen2, ADF, Synapse, Azure DevOps, Event Hubs, Key Vault Google Cloud Platform: GCS, BigQuery, Dataflow, Pub/Sub Data Engineering ETL/ELT Data Warehousing Dimensional Modeling Data Lake/Lakehouse Medallion Architecture CDC Data Quality Data Governance Metadata and Data Lineage DevOps / CI-CD Git Azure DevOps / GitHub Jenkins Terraform CI/CD automation Infrastructure as Code

Preferred / Nice-to-Have SkillsMLflow Databricks Machine Learning Feature Store Mosaic AI / GenAI dbt Apache Airflow Power BI / Tableau Delta Sharing Lakehouse Federation Liquid Clustering Data security and PII masking MLflow is particularly useful if the role touches ML/AI, as Databricks supports model tracking, lifecycle management, and deployment workflows alongside governed data.

QualificationsBachelor's degree in Computer Science, Engineering, Information Technology, or related field. 10+ years of experience in Data Engineering / Big Data / Analytics. 4+ years of hands-on Databricks experience preferred. Strong experience designing enterprise-scale data platforms. Demonstrated experience leading technical projects and mentoring engineers. Strong communication and stakeholder-management skills.

  • For applications and inquiries, contact:hirings@openkyber.com
Responsibilities

Loading...