Databricks Module Lead at Sopra Steria
Noida, Uttar Pradesh, India -
Full Time


Start Date

Immediate

Expiry Date

05 Oct, 26

Salary

0.0

Posted On

07 Jul, 26

Experience

5 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Databricks, PySpark, SQL, Python, ETL/ELT Pipelines, dbt, Apache Airflow, Azure DevOps, Git, CI/CD, Data Warehousing, Structured Streaming

Industry

Information Technology & Services

Description
Company Description About Sopra Steria Sopra Steria, a major Tech player in Europe with 51,000 employees in nearly 30 countries, is recognised for its consulting, digital services and solutions. It helps its clients drive their digital transformation and obtain tangible and sustainable benefits. The Group provides end-to-end solutions to make large companies and organisations more competitive by combining in-depth knowledge of a wide range of business sectors and innovative technologies with a collaborative approach. Sopra Steria places people at the heart of everything it does and is committed to putting digital to work for its clients in order to build a positive future for all. In 2025, the Group generated revenues of €5.6 billion. The world is how we shape it. Job Description Experience: 4 - 6 Years Location: Noida Education: B.E./ B.Tech. / MCA Must Have: Strong hands-on experience with Databricks for building and managing data engineering solutions. Expertise in PySpark for large-scale data processing and transformation. Advanced SQL skills, including complex queries, joins, aggregations, and performance tuning. Proficiency in Python for data processing, automation, and reusable code development. Experience developing and optimizing ETL/ELT pipelines for batch and near real-time data workloads. Ability to troubleshoot, monitor, and support production data pipelines. Good to Have: Experience with dbt for data modeling, testing, snapshots, and documentation. Knowledge of Apache Airflow for workflow orchestration and scheduling. Experience with Azure DevOps, Git, CI/CD pipelines, and deployment automation. Familiarity with data warehousing concepts including staging, intermediate, and mart layers. Understanding of data lineage, governance, and monitoring best practices. Exposure to Structured Streaming and real-time data processing. Qualifications Open Additional Information At our organization, we are committed to fighting against all forms of discrimination. We foster a work environment that is inclusive and respectful of all differences. All of our positions are open to people with disabilities.
Responsibilities
Lead the building and management of data engineering solutions using Databricks and PySpark. Develop and optimize ETL/ELT pipelines for batch and near real-time data workloads while supporting production pipelines.
Loading...