Designing and constructing high-performance, scalable data pipelines to operationalize our state-of-the-art solutions ranging from quantum computing pipelines to deep learning based recommender systems
Creating innovative solutions to tackle computationally intensive challenges and business issues
Writing production Python/Scala scripts in a Spark/Databricks environment to support the data pipelines that feed Machine Learning models
Leveraging cloud environments and following best practices while building solutions in the data lake as it evolves to meet the needs of the business
Performing data wrangling, feature engineering, and dimension reduction with data of all shapes and sizes, from parquet files to images
Partnering with Machine Learning Engineers, Data Scientists, and Cloud Architects to deliver models into production
Coaching, mentoring, and providing feedback to team members and junior data engineers
Discovering novel datasets and data sources and making them available for modeling and reporting
Using your expertise in big data technologies to champion innovation, design, and implement data-stores and analytic models, oversee large-scale data lake platforms, and design and implement efficient data pipelines from source systems to data-stores
Defining and implementing Quality Assurance best practices that support the implementation of data transformations and the availability of high-quality data.