Designed and operated global game-server infrastructure, cloud-provider automation, CI/CD systems, observability platforms, PostgreSQL high availability, DNS infrastructure, and registry automation.
KubernetesCI/CDObservabilityInfrastructure automationIncident responsePlatform interfacesTerraformGrafanaProduction systems
open to DevOps consultation through DevOps4Hire
Gabriel Boucher builds the systems behind the launch.
Platform / DevOps Engineer turning infrastructure, observability, and release work into systems teams can trust when production is awake.
- name
- Gabriel Boucher
- role
- Platform / DevOps Engineer
- mode
- systems built over time
- source
- github.com/gaboucher
launch.board runtime proof
provision ready deploy shipping observe watching Professional summary
resume.summary DevOps / Platform Engineer with 6+ years of experience designing and operating distributed cloud infrastructure.
- Designed and operated distributed cloud infrastructure with Kubernetes, Docker, GCP, AWS, Terraform, Go, and Python.
- Focused on automation, scalability, observability, and infrastructure cost optimization for production systems.
- Built cloud-provider abstractions, CI/CD automation, monitoring rules, image distribution systems, APIs, and internal services.
Experience
Edgegap / systems built over time database failover hardened
release/2026 deployed capability PostgreSQL high-availability architecture and migration
Role Lead engineer for the architecture and migration to a self-hosted highly available PostgreSQL platform.
Outcome Moved PostgreSQL from a single standalone instance to a 3-node highly available cluster with replication, pooling, routing, backups, and automatic failover.
- Designed the PostgreSQL high-availability architecture and migration strategy.
- Tested migration procedures across multiple infrastructure providers before production cutover.
- Implemented point-in-time recovery with pgBackRest and logical backups with pg_dump, storing backups in S3 buckets.
- Deployed and load-tested PgBouncer, HAProxy, and etcd for reliable connection pooling, routing, and automatic failover.
- Investigated replication failures, including multixact replay errors and replica recovery procedures.
signal costs reduced
release/2025 deployed capability Observability and monitoring cost optimization
Role Maintained monitoring, alerting, and metrics infrastructure for distributed production systems.
Outcome Reduced monitoring costs by about $2,000/month by analyzing metric cardinality and eliminating unnecessary ingestion while improving visibility.
- Built monitoring and alerting rules using Prometheus, Grafana, and PagerDuty.
- Analyzed cardinality of metric labels and applied fixes to reduce ingestion costs.
- Deployed a federated Prometheus metric collection system.
- Designed custom Grafana dashboards for client game-server metrics.
- Optimized container resource usage across monitoring APIs, proxies, game servers, and edge locations.
deploy path accelerated
release/2024 deployed capability Delivery reliability and DNS platform redesign
Role Built deployment pipelines and led the redesign of infrastructure DNS for high availability.
Outcome Reduced pipeline execution time by 5 minutes per commit, cut production configuration changes to under 1 minute, and reduced Redis CPU usage by about 95%.
- Developed unit and integration test pipelines using Bitbucket Pipelines and GitHub Actions.
- Reduced pipeline execution time by sharding tests for better efficiency and atomicity.
- Deployed an AWX stack for running Ansible workers and protecting environment secrets.
- Replaced a legacy DNS server with a replicated CoreDNS setup using Redis and automatic scaling.
- Modified the CoreDNS Redis plugin to use connection pooling, cutting Redis CPU usage from 100% to about 5% of 1 vCPU.
edge fleet expanded
release/2023 deployed capability Global edge infrastructure and gameserver orchestration
Role Designed and operated distributed game-server infrastructure across multiple cloud and edge providers worldwide.
Outcome Scaled game-server infrastructure to 615 locations while cutting orchestration overhead by about 40%.
- Built a Go API around nerdctl and containerd to orchestrate and monitor game-server deployments.
- Developed 17+ Python cloud-provider integrations to start virtual machines across 615 locations.
- Provisioned virtual machines, bare-metal servers, VPCs, disks, and supporting provider resources.
- Created a globally distributed DDoS protection service for TCP and UDP flood attacks.
- Deployed NATS to coordinate operating system updates during low-traffic periods.
image delivery hardened
release/2022 deployed capability Container platform and registry operations
Role Maintained and automated the multi-tenant container registry platform while modernizing deployment nodes.
Outcome Improved registry automation and storage management, supported uninterrupted image distribution, and cut 10 GB image pull times from about 1 minute to about 30 seconds from non-cached locations.
- Maintained a multi-tenant Harbor registry backed by S3-compatible storage.
- Managed Harbor upgrades using Helm without downtime.
- Built Python tooling around the Harbor API for clients, quotas, image layers, and garbage collection.
- Migrated Harbor PostgreSQL to a newer version without downtime.
- Migrated standalone-node deployments to K3s, a lightweight Kubernetes distribution suited to Edgegap infrastructure.
platform signals online
release/2021 deployed capability Monitoring foundation and game-server orchestration
Role Helped build the early operational foundation for game-server orchestration and production visibility.
Outcome Gave the platform its first dependable monitoring layer and helped shape the orchestration system that later supported global edge scale.
- Implemented monitoring and dashboards using Prometheus and Grafana.
- Helped design and develop the game-server orchestration system.
- Focused the early platform around automation, visibility, and production operability.
Earlier experience
LaPlank / operations under pressure - Prepared and cooked meals in a fast-paced kitchen.
- Managed ingredient inventory and supported team efficiency during peak hours.
- Followed sanitation standards and maintained cooking equipment.
Skills
stack.yml tracked / explainable / replaceable
platform map Cloud & infra
9 tools
- Kubernetes
- Docker
- Helm
- GCP
- AWS
- Linode
- Terraform
- Ansible
- Linux / Unix
Observe
6 tools
- Prometheus
- Grafana
- PagerDuty
- SQL
- PromQL
- MQL
Programming
8 tools
- Go
- Python
- Bash
- C#
- TypeScript
- Flask
- FastAPI
- Gin
Languages
3 tools
- French - Native
- English - Fluent
- Japanese - Beginner
Work with Gabriel through DevOps4Hire.
For platform engineering, DevOps automation, observability, scaling, or infrastructure cost work, the consultation starts on the main DevOps4Hire site.
Good fit for
Book a consultation - Kubernetes and cloud platform work
- CI/CD, provisioning, and automation
- Monitoring, incidents, and cost reduction