Senior Site Reliability Engineer - Midnight at IO Global

Remote, Scotland, United Kingdom -

Full Time

Start Date

Immediate

Expiry Date

13 Sep, 25

Salary

0.0

Posted On

15 Jun, 25

Experience

7 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Skills

Workable Solutions, Automation, Infrastructure

Industry

Information Technology/IT

Description

WHO ARE WE?

IOG, is a technology company focused on Blockchain research and development. We are renowned for our scientific approach to blockchain development, emphasizing peer-reviewed research and formal methods to ensure security, scalability, and sustainability. Our projects include decentralized finance (DeFi), governance, and identity management, aiming to advance the capabilities and adoption of blockchain technology globally.
We invest in the unknown, applying our curiosity and desire for positive change to everything we do. By fueling creativity, innovation, and progress within our teams, our products and services are designed for people to be fearless, to be changemakers.

Responsibilities

As a Senior SRE, you will be a key player in shaping the reliability and performance of our systems across our cloud infrastructure. You will design and implement solutions that improve our service reliability, automate routine tasks, and facilitate smooth collaboration between development and operations teams. This role demands a blend of deep technical expertise, a proactive mindset, and the ability to take vague or evolving challenges and refine them into robust, workable solutions.

Infrastructure & Automation:
- Design, build, and maintain scalable and highly available systems, primarily on AWS, using best practices.
Manage and optimize Kubernetes clusters for high availability and performance, extending them when it makes sense to expand functionality.
Leverage GitOps principles to automate deployments and manage container orchestration.
Implement and manage CI/CD pipelines ensuring seamless, high-quality deployments, finding and removing bottlenecks, improving performance and working alongside teams to refine feedback loops and automate toil away.
Develop automation tools and scripts to improve operational efficiency.
Monitoring & Incident Response:
Implement robust monitoring solutions with Prometheus and related tooling to ensure system health and performance.

Participate in on-call rotations and lead incident response efforts, turning challenges into learning opportunities.
Collaborate with dev teams to define and implement SLOs/SLIs
Problem Solving & Communication:
Take vague or loosely defined problems, work closely with cross-functional teams, and distill them into clear, actionable plans.

Communicate technical solutions and incident retrospectives effectively across both technical and non-technical stakeholders.
Innovation & Continuous Improvement:
Evaluate and adopt new technologies, with a special advantage for candidates with blockchain experience, to keep our systems at the cutting edge.

Document processes and best practices, ensuring that knowledge is shared across the team and continuously improved.
Strive to strike a balance between effective delivery of goals and a measurable high standard of these goals. Always apply a layer of polish and due diligence when delivering.