Develop and optimize YDB components to take full advantage of modern hardware, including QLC NVMe drives, high-bandwidth network adapters, and DPUs.
Improve performance across widely used storage technologies, including HDDs and TLC NVMe devices.
Reengineer system components using more efficient algorithms to address complex scalability, performance, and reliability challenges.
Design and implement high-performance, low-latency software components for heavily loaded production systems.
Analyze system behavior and identify performance bottlenecks using profiling, debugging, and diagnostic techniques.
Work on distributed storage and database infrastructure supporting demanding AI and cloud workloads.
Contribute to the reliability, availability, and durability of production-grade distributed systems.
Participate actively in incident investigation and resolution, helping identify root causes and implement lasting improvements.
Collaborate with engineers across infrastructure, storage, networking, and AI-focused teams to deliver robust technical solutions.
Contribute to an open-source codebase and participate in technical discussions, design decisions, and continuous improvement.
Take ownership of engineering challenges from investigation and design through implementation, optimization, testing, and production deployment.
Requirements
5+ years of professional experience developing in C or C++ for highly loaded, performance-critical systems.
Strong understanding of low-level systems concepts, including CPU caches, modern CPU atomic operations, and NUMA architectures.
Proven experience developing high-performance and low-latency software components.
Ability to investigate complex production issues using core dumps, flame graphs, sanitized builds, and related debugging techniques.
Strong systems programming and performance optimization skills, with an analytical approach to diagnosing bottlenecks.
Experience with on-disk data structures such as LSM trees or B+ trees is a strong advantage.
Hands-on experience with tools such as perf, VTune, bpftrace, or gdb is desirable.
Knowledge of storage algorithms, including erasure coding and checksumming, is a plus.
Understanding of storage device internals, particularly NVMe and HDD technologies.
Familiarity with Linux kernel development and technologies such as SPDK, DPDK, libaio, or io_uring is advantageous.
Knowledge of networking concepts and protocols including IP, TCP, UDP, and DNS; experience with InfiniBand, RoCE, or RDMA is a plus.
Experience with Kubernetes and Grafana is desirable.
Experience designing production-grade distributed storage components and understanding availability and durability calculations is highly valued.
Demonstrated ability to troubleshoot complex systems, participate in incident resolution, and work effectively in collaborative engineering environments.
Strong ownership, problem-solving skills, and enthusiasm for tackling technically challenging infrastructure problems.
Willingness to participate in coding interviews as part of the recruitment process.