Production Database Operations: Own the day-to-day operation, health, availability and continuous improvement of Toku's production MySQL databases running on AWS RDS, ensuring reliability across our global SaaS platform.
Incident Management & Operational Ownership: Take ownership of production incidents, perform troubleshooting and root cause analysis, restore service quickly, and drive permanent improvements that reduce repeat operational issues.
Automation & Process Improvement: Identify repetitive operational work, evaluate where automation delivers the greatest value, and build scalable workflows and tooling that reduce manual effort, operational risk and engineering overhead.
Database Workflow Engineering: Improve how database changes, production data fixes and operational requests are managed by introducing safer, more auditable and scalable engineering processes.
Engineering Collaboration: Partner closely with software engineers to understand application behaviour, investigate production issues, interpret logs, and improve the interaction between engineering and database operations.
Database Performance & Optimisation: Monitor and optimise database performance through query tuning, indexing strategies, capacity planning and proactive performance analysis.
AWS Database Engineering: Manage and continuously improve AWS database services, with a strong focus on Amazon RDS while supporting Toku's ongoing evolution towards Amazon Aurora.
Monitoring & Observability: Utilise monitoring platforms such as Datadog and AWS Performance Insights to identify trends, troubleshoot production issues, and improve operational visibility.
Security, Compliance & Governance: Apply secure database practices, improve operational governance, support audit readiness, and help reduce operational risks associated with production database management.
High Availability & Disaster Recovery: Support and continuously improve high availability and disaster recovery capabilities while leveraging modern AWS managed database services.
Continuous Improvement: Challenge existing ways of working, identify opportunities to eliminate operational toil, and contribute ideas that improve the long-term scalability of the database platform.
On-call Support: Participate in a shared on-call rotation (approximately two weeks per month), responding to production incidents outside normal business hours when required.