
Resume
Table of Contents
Samson Gama#
Senior SRE / DevOps Engineer | Kubernetes | Go | Cloud Infrastructure | OSCP
Vancouver, BC
Summary#
Senior SRE/Platform Engineer with 8+ years of experience designing, building, and operating distributed systems, Kubernetes platforms, cloud infrastructure, and high-scale production services. Experienced in reliability engineering, incident response, performance engineering, infrastructure automation, and observability, with recent experience operating GPU infrastructure for AI/HPC workloads. Strong background in Go, Kubernetes, Terraform, GitOps, and production systems at scale.
Professional Experience#
Bitcomplete#
Senior Service Reliability Engineer, GPU Infrastructure | Mar 2026 - Jun 2026
- Improved operational readiness for GPU infrastructure, establishing standardized runbooks, alerting, and incident-response workflows to improve reliability and accelerate incident resolution across AI/HPC workloads
- Designed and executed the consolidation of fragmented ArgoCD-managed GPU environments, eliminating deployment drift and standardizing NVIDIA tooling across clusters, reducing customer workload downtime
Demonware#
Senior Service Reliability Engineer | Aug 2020 - Jan 2026
- Applied SRE practices to improve reliability, scalability, and operational efficiency across production systems, establishing company-wide conf.d configuration standards and improving resource utilization
- Designed, built, and owned a distributed load-testing platform used by engineering teams company-wide to validate services against large-scale traffic spikes, using Kubernetes, Go, Redis, and distributed rate limiting
- Diagnosed and resolved complex production incidents across distributed systems during peak live traffic, developing runbooks and operational improvements that reduced repeat incidents and alert fatigue
- Partnered with product and platform engineering teams to improve service architectures through pull requests, code reviews, technical design, technology selection, RFCs, and load-testing infrastructure
- Evaluated and helped drive adoption of Vitess for MySQL workloads on Kubernetes, developing PoCs, design documents, RFCs, and migrations to reduce the overhead of maintaining MySQL clusters on VMs
- Mentored engineers and promoted service ownership and operational excellence, helping teams adopt more reliable engineering and production practices by taking pride in our work for the video game series Call of Duty
Mastercard (Nudata Security)#
Senior Software Engineer (DevOps) | Oct 2018 - Jul 2020
- Designed and developed highly available microservice infrastructure across AWS regions using CloudFormation and SaltStack, including backup strategies, disaster recovery procedures, and operational runbooks
- Implemented transparent production traffic mirroring to reproduce and debug live production issues safely, improving root-cause analysis and reducing time to resolution
- Developed and enhanced operational tooling and observability to improve logging, error handling, and notification workflows for background jobs, reducing manual developer intervention and improving visibility into failures
- Led bi-weekly technical sessions on offensive security, covering practical attack techniques and security concepts to improve engineering security awareness and defensive practices
Absolute Software#
Software Developer | Jun 2017 - Sep 2018
- Developed and executed MongoDB data migrations that improved production performance by 20% while maintaining application compatibility
- Re-engineered a microservice from Python to Go, achieving a 20x increase in throughput and 90% reduction in CPU and memory utilization while maintaining integration test coverage
- Built internal testing and simulation tooling to replicate and simulate hundreds of thousands of devices in development environments, enabling realistic scalability and performance testing
- Collaborated with engineers on Docker, Python, Go, Kubernetes, and infrastructure best practices, improving consistency across development and production environments
- Automated repetitive deployment workflows through Jenkins, reducing manual operational work and improving engineering productivity
Prizm Media Inc#
Software Developer Intern | Dec 2015 - Aug 2016
- Developed Android and iOS applications for an AWS-hosted fitness-oriented social networking platform, contributing across mobile application development, and backend services
- Developed and evaluated machine-learning solutions using support vector machines (SVMs) and neural networks to address application-specific problems, applying model development and performance evaluation techniques
- Reduced response times for frequently used APIs by up to 90% through database query optimization and application-level caching, improving application responsiveness and reducing backend processing overhead
Ericsson#
Software Developer Intern | Sep 2015 - Dec 2015
- Developed an OpenStack Neutron plugin in Python to automate lifecycle management of virtual Ericsson routers, integrating network infrastructure with OpenStack and improving infrastructure provisioning consistency
- Troubleshot containerized virtual-router environments and automated test pipelines on Mirantis OpenStack, diagnosing infrastructure, networking, and deployment issues to improve reliability and test consistency
Grin Technologies#
Software Developer Intern | May 2014 - Aug 2014
- Developed two web applications that processed and visualized raw electric-vehicle telemetry data, enabling users to plot trip data and generate customized wheel configurations using dynamic frontend and backend data
- Administered an electric-vehicle forum serving 5,000+ daily users on AWS through 2020 - endless-sphere.com
Technical Skills#
| Category | Skills |
|---|---|
| SRE & Reliability | Reliability Engineering, Distributed Systems, Scalability, Incident Response, Root Cause Analysis, Monitoring, Alerting, Runbooks, Disaster Recovery, Performance Engineering, Load Testing |
| Observability | Prometheus, Grafana, StatsD, Graphite, VictoriaMetrics, OpenTelemetry |
| Programming | Go, Python, Bash, C/C++, Java, JavaScript, Rust |
| Cloud & Infrastructure | AWS, GCP, Azure, Kubernetes, Terraform, Ansible, Packer, Helm, Linux |
| CI/CD & GitOps | ArgoCD, Jenkins, GitOps, CloudFormation, SaltStack |
| Databases & Distributed Data | MySQL, PostgreSQL, MongoDB, Redis, Cassandra, Kafka, Vitess |
| Networking | Cilium, Nginx, HAProxy, OpenResty, MikroTik |
| Security | Linux Hardening, Offensive Security, OSCP, CEH |
Projects#
Self-Hosted Homelab & Cloud Infrastructure | 2010 - Present#
- Built and managed multi-cloud infrastructure across GCP, DigitalOcean, and Cloudflare using Terraform, Ansible, Packer, and GitOps, with Prometheus and Grafana observability across Linux workloads and infrastructure
- Designed and operated a self-hosted Kubernetes platform on Proxmox, using Talos Linux, Cilium CNI, ArgoCD, Helm, Terraform, and GitOps to provide reproducible infrastructure and automated application deployments
- Engineered a production-grade homelab environment spanning MikroTik VLAN/firewall networking, custom OpenResty/Lua load balancing, authentication, and hardened Ubuntu infrastructure
- Implemented disaster-recovery and S3 backup/restore testing and deployed honeypots to validate recovery procedures and detect malicious activity
Leadership and Community#
Strata Council President#
- Lead a council representing a community of ~1,200 residents, coordinating stakeholders, facilitating meetings and decision-making, and communicating community priorities and operational issues
Veazey Foundation Trustee#
- Serve on the board of a charitable foundation providing full scholarships to incoming UBC students, contributing to governance, financial decisions, and long-term planning
Achievements#
- Trend Micro CTF 2019 Global Finalist | Maple Bacon, UBC CTF | 12th of 1,000 teams | 2019
Education#
- University of British Columbia, B.A.Sc., Computer Engineering | 2017
There are no articles to list here yet.