Distributed Systems · Technical Leadership

Turning ambiguous production problems into
boringly reliable and efficient systems.

Distributed systems engineer focused on reliability and operability—from embedded runtimes and distributed storage to high-throughput, real-time processing pipelines and control planes. I work across teams on system design, production migrations, and the tooling needed to operate those systems safely.

Get in touch Download CV

Define the behavior, design for failure, validate the change, and leave the system easier to operate

01

I make the decision explicit: the production constraint, expected behavior, technical trade-offs, cross-team dependencies and ownership, and evidence that will validate the approach.

02

I carry the change through end-to-end tests, live-traffic validation where appropriate, staged rollout, tested recovery procedures, and operational handoff.

03

I turn recurring operational work that requires expert intervention into tooling and validated, automated self-service workflows so other engineers can operate and extend the system independently.

Where this approach helps most

I help teams make high-throughput production systems more reliable, diagnosable, and easier to operate—from processing pipelines, routing and control planes to distributed storage.


Selected Evidence

Cross-domain monitor correctness during partial degradation

Worked across Alerting and Metrics to make monitor evaluation aware of ingestion-to-query health, reducing the risk that delayed or incomplete platform data would be interpreted as customer behavior; the framework supported safer degraded-mode recovery and was extended to additional product domains.

High-throughput realtime processing with ~3× lower resource use

Owned and optimized routing and processing systems handling billions of events per second across petabyte-scale data and thousands of workers, reducing resource use by ~3× through caching, sharding, and hot-path optimization while validating correctness and failure behavior under production conditions.

Distributed storage and synchronization

Stabilized correctness-sensitive distributed filesystem and synchronization internals, then productized their Linux operation so users could diagnose and recover runtime behavior.

Embedded platform reliability

Owned firmware, kernel, boot, and recovery work that made embedded wireless platforms safer to update, diagnose, and operate remotely.


Career History

Dec 2017 - Present

Datadog · Paris, France

Distributed Systems Software Engineer

Built and operated routing control planes and high-throughput processing systems for Datadog's Alerting and Metrics products. Worked with Product and platform teams on monitor behavior, telemetry diagnostics, and production migrations.

  • Authored the design and led the evolution of a routing and migration control plane that enabled self-service workflows, including AI-assisted execution and troubleshooting. Protected production migrations with staged rollouts, live-traffic validation, and simple recovery procedures.
  • Defined customer-facing monitor semantics and operator-facing telemetry diagnostics with Product, engineering leadership, and platform teams. Preserved compatibility for established monitor workflows while clarifying recommended evaluation behavior. Redirected a high-cardinality telemetry investigation proposal from a costly standalone pipeline to a bounded investigation workflow, and led its design and production delivery.
  • Owned the integration of data-freshness and completeness signals across Metrics processing and query systems and the Alerting evaluation pipeline, reducing false positives and false negatives during partial degradation and making degraded-mode recovery more predictable.
  • Improved service-check and realtime metrics processing pipelines that handled billions of events per second, processed petabyte-scale datasets, and ran on thousands of worker pods. Cut resource use by approximately 3x through caching, sharding, and hot-path optimization.
  • Built reusable migration tooling, AI-assisted troubleshooting, and end-to-end validation for the realtime storage tier. Engineers outside the originating team used the workflow to complete migrations and cleanup without direct support.
Apr 2015 - Dec 2017

Lima · Paris, France

Core Software Engineer

Built core components of a distributed filesystem, synchronization subsystem and manufacturing tooling.

  • Improved replication robustness, transaction performance, and garbage-collection efficiency so the distributed filesystem could handle very large trees with hundreds of thousands of files and directories.
  • Designed the factory test and provisioning process, built the software for hardware testing, firmware flashing, and device identification, and installed the system on-site.
  • Designed and implemented file sharing via public web gateways.
Dec 2008 - Apr 2015

Sequans Communications · Paris, France

Embedded Software Engineer

Developed LTE firmware and embedded platform software, including Linux and real-time kernels, boot and recovery, device drivers, and build infrastructure.

  • Implemented JFFS2 for vxWorks in read/write mode to simplify transition from vxWorks to eCos kernel.
  • Optimized USB, networking, and inter-processor layers, bringing throughput closer to hardware limits.
  • Made boot, recovery, and firmware updates more resilient. Built tools for post-crash diagnosis and reproducible multi-repo firmware builds.

Expertise

High-Throughput Stream Processing Reliable Control Planes and Routing Distributed Storage and Synchronization Embedded Systems, Kernels, and Drivers Production Reliability and Performance

Let's work together

Dmytro Milinevskyi

Distributed Systems · Technical Leadership · Paris, France or Remote