MIGOps
Safe NVIDIA MIG operations with workload-aware checks, profile discovery, planning, snapshots, drift detection, and dry-run/apply workflows.
$ migops status $ migops recommend gpu 0 4 $ sudo migops split gpu 0 4 --yesopen repository ↗
AI Infrastructure & Systems Engineer
AI infrastructure, system administration, virtualization, cloud, AWS, NVIDIA GPU infrastructure, Linux, Kubernetes, and automation.
Public tools built around NVIDIA GPU infrastructure and day-to-day operations.
Safe NVIDIA MIG operations with workload-aware checks, profile discovery, planning, snapshots, drift detection, and dry-run/apply workflows.
$ migops status $ migops recommend gpu 0 4 $ sudo migops split gpu 0 4 --yesopen repository ↗
GPU node diagnostics across telemetry, PCIe, ECC/Xid, NVLink, Fabric Manager, DCGM, CUDA, containers, MIG, and Kubernetes GPU resources.
$ gdiag $ gdiag stack $ gdiag report -o gpu-report.htmlopen repository ↗
My main areas of hands-on infrastructure work.
AI compute platforms, GPU servers, cluster infrastructure, deployment, operations and troubleshooting.
Linux administration, systemd, services, logs, RAID, LVM, storage, performance and recovery.
VMware vSphere, vCenter, ESXi, Proxmox, virtual machines, GPU passthrough, OpenStack and OpenNebula.
AWS, EC2, IAM, cloud infrastructure, access control and compute operations.
NVIDIA GPUs, CUDA, MIG, NVML, DCGM, Fabric Manager, drivers and GPU diagnostics.
Ubuntu, server administration, shell, networking, storage, troubleshooting and automation.
Kubernetes, K3s, Docker, Rancher, workloads, GPU resources and cluster operations.
Terraform, Python, Bash, Git, YAML, scripting and repeatable operational workflows.
TCP/IP, VLANs, 5G, QoS, latency testing, connectivity and edge infrastructure.
Dell PowerEdge, HPE servers, GPU hosts, storage, hardware validation and lifecycle operations.
Most of my work is hands-on: deploying systems, troubleshooting issues, maintaining infrastructure, and automating repetitive tasks.
My main focus is AI infrastructure, Linux system administration, virtualization, cloud platforms, AWS, and NVIDIA GPU infrastructure, with Kubernetes and automation around those environments.
Industry certifications across AI infrastructure, virtualization and cloud.
February 2026
February 2023
October 2022
September 2019