Cloud Platform DevOps Engineer - Assistant Vice President at Citi canada
ontario, Ontario, Canada -
Full Time


Start Date

Immediate

Expiry Date

29 Nov, 26

Salary

0.0

Posted On

31 Aug, 26

Experience

5 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Yes

Skills

Industry

Information Services

Description
  • Design & Implementation: Lead the design, implementation, and ongoing management of secure, scalable, and resilient infrastructure components.
  • Secret & Certificate Management: Administer and maintain secret and certificate management solutions using HashiCorp Vault, including policy definition and integration.
  • Database Management: Perform hands-on administration and optimization of database systems (PostgreSQL, Oracle, MongoDB), including performance tuning, backup, and recovery strategies.
  • Workflow Orchestration: Deploy, monitor, and troubleshoot data orchestration workflows using Apache Airflow, and develop/optimize DAGs.
  • Messaging Systems: Implement and manage messaging queues such as Kafka and IBM MQ, including cluster setup and configuration.
  • API Integrations: Develop, maintain, and troubleshoot RESTful API and SOAP integrations critical for system connectivity.
  • Build Automation: Implement and optimize build and deployment processes using Gradle.
  • Container Orchestration: Design, implement, and manage container orchestration platforms with Kubernetes and Helm, including integration with CyberArk and HashiCorp for secrets management. Create, debug, and troubleshoot Kubernetes PODs, Jobs, and Deployments using YAML.
  • Storage Management: Configure and manage persistent storage solutions including PVC, SONiC NAS, and S3, with an awareness of storage requirements for AI/ML workloads.
  • Networking & Load Balancing: Set up and maintain load balancing solutions (e.g., Nginx, HAProxy, AWS ELB/ALB, Kubernetes Ingress controllers) for high availability and performance.
  • Monitoring & Logging: Implement, configure, and utilize comprehensive monitoring and logging solutions (Prometheus, Grafana, ELK Stack) to ensure system health and proactively identify issues, including those relevant to AI/ML applications.


How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities
Loading...