Data Engineer at HCLTech
Ontario, Ontario, Canada -
Full Time


Start Date

Immediate

Expiry Date

29 Dec, 26

Salary

45000.0

Posted On

30 Sep, 26

Experience

2 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Yes

Skills

Industry

Consumer Services

Description

Job SummaryWork with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.•Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.•Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.•Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns•Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.•Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and featuresIdeal Candidate Qualifications:•Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).•High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.•Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)•Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.•Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.

How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities

Job SummaryWork with cutting-edge big data platforms (e.g., Databricks, Apache Spark) at large scale, pushing the boundaries of data processing and model enablement.•Build and maintain robust ETL/ELT pipelines for ingestion, transformation, and aggregation of large-scale datasets on Hadoop and enterprise data platforms.•Develop high-performance data processing jobs using PySpark/Spark, Python on data platforms such as cloudera and databricks.•Optimize pipeline performance and cost through partitioning, file formats, compute tuning, and efficient query patterns•Contribute to CI/CD for data workflows (testing, code reviews, deployment automation), promoting engineering best practices and maintainable codebases.•Partner with Product Managers to develop a deep understanding of users and use cases and apply that knowledge to scoping and building new modules and featuresIdeal Candidate Qualifications:•Strong hands-on experience in data engineering building production-grade pipelines on big data platforms (Hadoop ecosystem and cloud data platforms - databricks).•High proficiency in using Python, Spark, Hadoop platforms & tools (Hive, Impala, Airflow, NiFi), SQL to build Big Data products.•Hands-on experience with cloud data platforms such as databricks, snowflake (databricks preferred)•Experience with orchestration/integration tools such as Apache Airflow, Apache NiFi, or Talend.•Working knowledge of DevOps/CI-CD practices: version control (Git), automated testing, release pipelines, and observability.

Loading...