I recently started getting hands-on with GPU infrastructure. Before running any workloads on the cluster, I wanted to benchmark the network and see what the hardware was actually capable of. The setup was two servers with eight NVIDIA H200s in each, so sixteen GPUs. I’ve spent about twenty years in ordinary IT infrastructure, but I’m […]
Category: AI
Beyond the Dashboard: Why the Future of SRE is Conversational
We’ve all been there: an alert fires at 2:00 AM. In the old days, you’d manually grep logs across half a dozen systems. Today, modern observability tools are already very good at connecting the dots – using automated root cause analysis to tell us which microservice caused a latency spike. But connecting the dots isn’t […]
