Role Overview
We are looking for an experienced Data Engineer with strong expertise in PySpark, Python, SQL, Databricks, and cloud data platforms to design, develop, and maintain scalable data pipelines and data processing solutions.
The ideal candidate should have hands-on experience working with large datasets, developing ETL/ELT pipelines, implementing data transformations, optimizing Spark workloads, and building cloud-based data platforms.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows.
- Develop high-performance data processing solutions using PySpark and Apache Spark.
- Build data pipelines to ingest data from databases, APIs, files, applications, and other sources.
- Perform data cleansing, transformation, aggregation, and validation.
- Develop and optimize Spark SQL and PySpark jobs.
- Work with structured, semi-structured, and unstructured data.
- Implement data pipelines using Databricks and cloud-based data platforms.
- Optimize Spark jobs for performance, scalability, memory utilization, and processing time.
- Implement data quality, validation, monitoring, and error-handling mechanisms.
- Work with data lakes, lakehouses, and data warehouse environments.
- Develop reusable and scalable data engineering frameworks.
- Collaborate with Data Architects, Data Scientists, Business Analysts, and application teams.
- Troubleshoot production data pipeline issues and perform root-cause analysis.
- Participate in code reviews, testing, deployment, and production support.
- Maintain technical documentation for data pipelines and data architecture.