Work sent me to Google Cloud Next this year. Almost everything was about AI. There were even banners in the hallway saying “Put an agent on it”. The engineering talks, like the GKE one, ran in small rooms with thin crowds.
The GKE talk
My favorite talk opened with Yahoo Mail moving petabytes of data onto GCP (yes, they’re still around). They split GKE into separate clusters for ingress, mail, add-ons like contacts and calendar, and analytics, so a runaway batch job can’t take down mail delivery. They picked GCP for Spanner’s multi-region consistency and to run GKE next to that data.
Google spent the rest of the talk on autoscaling and capacity, aimed at staying fast without over-provisioning:
Predictive HPA forecasts demand from history and scales out a few minutes ahead, for workloads with long startups and cyclical traffic.
Native custom-metric HPA reads a Prometheus metric straight off the pod’s /metrics through an AutoscalingMetric CRD, skipping the Cloud Monitoring adapter and its pile of IAM setup. The demo scaled on vLLM’s KV-cache usage.
Compute classes let you hand GKE a prioritized list of machine types (spot first, on-demand fallback) and have it find capacity, with node-pool auto-creation you can scope per class instead of cluster-wide.
Capacity buffers keep spare nodes warm so critical workloads can scale out without waiting for provisioning. The speaker pitched them as an alternative to balloon pods, low-priority placeholder pods parked to hold space.
VPA can now resize a running pod in place instead of restarting it, so rightsizing doesn’t interrupt the service.
CPU Boost uses that in-place resize to give a pod extra CPU during startup and take it back once the pod is ready, so a slow-starting JVM doesn’t hold that headroom for life.
The CPU Boost config they showed:
startupBoost:cpu:type:"Factor"factor:3# 3x the CPU request and limit while startingdurationSeconds:10# keep the boost 10s past readiness
BigQuery
The BigQuery talk was mostly about making it a better source of context for models:
Fluid Scaling drops the 60-second minimum charge on autoscaled queries, so you pay per second for bursty, agent-driven traffic.
ObjectRef is a column type pointing at an object in storage (an image, an audio file), so you can join unstructured data with structured rows and run AI/ML on it in one query.
Hybrid Search runs vector and lexical search together and merges them with a reranker, for RAG.
Observability gained job-level cost attribution to answer “who is burning the budget”.
They also made it easier to call LLMs from SQL. An optimized mode for AI.CLASSIFY and AI.IF has Gemini label a sample of rows, then trains a small model on those labels to handle the rest. Google says that uses 230x fewer tokens than calling Gemini on every row. It still sounds like an easy way to accidentally burn a ton of money.
The AI talks
Gemini Enterprise was the keynote’s centerpiece and looked like Google’s Glean plus a drag-and-drop GUI for building no-code agent flows from predefined steps.
BigLake is now the Cross-cloud Lakehouse, built on Apache Iceberg so you can query data sitting in other clouds without copying it.
AlloyDB is Google’s Postgres-compatible answer to Aurora, and it got optimized modes for its AI functions too, like ai.if.
Google’s ADK agent framework picked up graph-based and multi-agent workflows and Agent Skills for Python, Go, and TypeScript.
Semantic layers like LookML have been around for years, but they matter more once an LLM writes the queries. A Looker talk said the descriptions in LookML help LLM accuracy a lot. It also seemed like a pitch for running more LLM calls through Looker, where the tokens bill separately.
L’Oreal’s talk was about opening agents to non-programmers on the business side. They built their agent platform on ADK, with MCP tools that can even pull reports, and the talk showed off some Gemini Enterprise too. I like the goal, since programmers shouldn’t be the only ones who get to build useful things. The speaker was vague about what business users are allowed to do beyond sticking to the guardrails and brand guidelines. Code review never came up, though an earlier slide seemed to show an approval process. I’d want some review, mainly around exfiltration. If those agents are limited to Gemini Enterprise’s predefined steps, how much one could leak depends on whether any step can make an arbitrary HTTP request. I couldn’t tell, and the controls I’d want might exist and just not have made the talk.
A separate Wayfair talk covered the Universal Commerce Protocol, a standard they’re building with Google and other retailers so LLMs can surface products and you can check out without leaving the chat. They said their LLM-referred traffic doubled in two quarters and converts better than their usual channels, though it’s still a single-digit share of total traffic.
The expo floor
Google had the accelerators behind AI Hypercomputer out on the floor. One was an Ironwood TPU rack, where a full pod is 9,216 liquid-cooled chips sharing 1.77 petabytes of HBM over an optical fabric that steers light between fibers with tiny tilting mirrors. The other was an A4X Max rack, NVIDIA’s GB300 NVL72, with 72 Blackwell Ultra GPUs pulling around 120 kW.
Boston Dynamics had Spot patrolling a mock industrial site, and Google and Mars ran an “AI Snack Factory” where Gemini invents candy flavors and attendees vote. “Golden Latte Flambe” was winning.
Macvlan is a Docker networking mode that gives each container its own MAC and IP on the network the host port is plugged into. The container shows up on the subnet as its own device, gets a DHCP lease from the router, and talks to the LAN directly instead of going through the Docker bridge.
I have a box with multiple Ethernet ports, each port on a different VLAN. Some containers sit on the port assigned to a guest-style VLAN that can only reach the internet, not other hosts on the LAN. That worked fine in the default bridge macvlan mode while each port only had one service on it. Once I added more services to the same port, I hit a problem I hadn’t thought about: containers sharing a macvlan parent in bridge mode can freely ARP and talk to each other inside the kernel, which defeats the VLAN-level isolation.
Macvlan has a private mode for exactly this. The kernel drops frames going between siblings on the same parent. Outbound traffic via the gateway still works. No switch cooperation, no extra VLANs.
One caveat: the host itself can’t reach its own containers over that interface. That’s true for any macvlan setup, not specific to private mode. For my case it’s fine since the containers just need internet access and to be reachable from other subnets through the router.
More on the modes underneath: ip-link(8) documents macvlan’s private, vepa, bridge, and passthru modes. Docker’s guide: Docker macvlan driver.
Distributing one-off Python scripts has gotten a lot better recently. PEP 723 added a standard way to declare dependencies and a minimum Python version inside a single .py file, and uv makes running them painless. I used this to share a script with the team and it just worked across everyone’s machines.
Before this, you’d end up writing some wrapper shell script that creates a virtualenv, makes sure the right Python version is being used, installs the dependencies, and tries to keep everything in sync. Brew-installed Python was especially bad for this because brew upgrade can swap your Python version out from under every virtualenv pointing at it. Even version managers like asdf or pyenv don’t really help because the recipient still needs the right version installed through their manager. A lot of ceremony for a single file.
The # /// script block is TOML. You don’t have to write it by hand either, uv add --script example.py 'requests' will add it for you.
The recipient just needs uv installed:
uv run check_status.py
uv reads the inline metadata and installs the dependencies in an isolated environment. If the required Python version isn’t on their machine, uv downloads it automatically. No virtualenv, no version manager, no wrapper script.
You can also lock dependencies with uv lock --script example.py for reproducible runs. For bigger tools that span multiple files, uv tool install is the next step up and can install from a private PyPI index.
I finally got to a local LLM setup that feels pretty usable within a 16 GB VRAM constraint (around 40 tok/s).
As far as I can tell, Open WebUI is still the best open source chat interface for local models. I am running it on a Ryzen 5 5600 box with an RTX 5060 Ti 16 GB card, with Qwen3.5-9B GGUF at Q8_0 as the main local model. I started with Ollama because it is the popular default. For Qwen 3.5 though, I hit a few open issues at the time that made llama.cpp easier for my setup: much slower inference than llama.cpp with the same model, long stalls on later turns in a conversation, and broken /no_think handling for Qwen. It is a very new model, and I am sure the Ollama project will get those fixed.
llama-server exposes an OpenAI-compatible API, so Open WebUI mostly treated it like a config swap.
The more interesting part is that Qwen 3.5’s newer hybrid architecture seems to help a lot here. The 9B model feels much better than I would have expected for the size, and this is the first local setup I have had on this machine that felt worth keeping around.
From what I had been reading, thinking mode also did not seem worth the wait for a setup like this, and that matched what I was seeing. Qwen 3.5 defaults to thinking, but for my use it was usually slower, often added 30-60 seconds, and sometimes got stuck in long thinking loops without noticeably improving the answer. LLAMA_ARG_REASONING=off (--reasoning off) turns it off, with LLAMA_ARG_JINJA=1 so llama.cpp uses the chat template from the GGUF.
Most of the other tuning was about using the remaining VRAM well once the 9B model itself had already taken roughly 9 GB. LLAMA_ARG_N_PARALLEL=1, LLAMA_ARG_FLASH_ATTN=1, and a q8_0 KV cache were what got the whole setup to around 14 GB and let me spend the rest on context instead of spilling into system memory. The next size up looked much more likely to spill into RAM. I wanted to see what I could get done with this Nvidia card first, since it still seemed like the safest compatibility bet.
For serious work I would still use frontier hosted models. The local setup is useful enough to keep around for narrower cases, especially redacting or cleaning text before sending it to a cloud model.
Full example gist with the commented minimal two-container compose file:
I recently got a small fix merged into metrics-server, which powers kubectl top and is used for horizontal pod autoscaling in Kubernetes. It’s not exactly a core component, but most production clusters have it installed.
I’ve been using ArgoCD at work lately for deploying Helm charts through a GitOps flow, and I noticed that the metrics-server APIService resource kept showing as “OutOfSync” even though nothing had actually changed.
The issue was that the Helm template always rendered the insecureSkipTLSVerify field explicitly, but Kubernetes omits it from live resources when it’s false (the API default). This caused ArgoCD to see a constant diff.
The fix was to conditionally render the field using {{- with .Values.apiService.insecureSkipTLSVerify }} so it only appears when set to true. Same approach other projects like KEDA have used.
It’s a tiny fix, but it’s satisfying to have a change merged into something as widely deployed as metrics-server.
Earlier this year, I wanted to help seed Linux distribution torrents using cheap VPS servers that offer terabytes of monthly bandwidth for less than $20/year. To maximize cost efficiency, I needed something as memory-efficient as possible.
After trying qBittorrent, rTorrent, and Transmission (via Docker) on a 1GB RAM VPS, I kept running into OOM issues and configuration headaches. Those clients are great, but running them in 1GB RAM doesn’t seem to be a design goal. It’s probably still possible with enough tweaking.
I ended up building distro-seed, a lightweight Go-based BitTorrent seeder using anacrolix/torrent. It’s a Go library that handles all the BitTorrent protocol details while letting you build exactly what you need without the overhead of a full client.
So far, it’s seeded 1.25 TB of Linux distributions.
The project includes an Ansible playbook to set up a fresh Ubuntu VPS for automatic seeding. The whole setup is simple: configure your torrent sources in a YAML file, run the playbook, and let it seed.
Today, I came across a bug where a Celery worker wasn’t gracefully shutting down, and it was causing some odd “Connection Refused” errors from requests within the task being run by the worker. It was also shutting down before it could send errors to Rollbar/Sentry for the team to know they need to address it.
This was happening because the entrypoint had a script that effectively did this:
echo "starting celery worker"
celery -A tasks worker
I was recently troubleshooting some packet loss with ping on linux, and I noticed by default ping won’t explicitly show lost packets:
$ ping x.x.x.x
64 bytes from x.x.x.x: icmp_seq=8 ttl=52 time=18.1 ms
64 bytes from x.x.x.x: icmp_seq=11 ttl=52 time=21.2 ms
(notice the skipped 8-10, those were lost packets)
To fix this, you can add run ping with -O:
$ ping -O x.x.x.x
64 bytes from x.x.x.x: icmp_seq=11 ttl=52 time=19.0 ms
no answer yet for icmp_seq=12
no answer yet for icmp_seq=13
64 bytes from x.x.x.x: icmp_seq=14 ttl=52 time=18.9 ms
To improve this output, try adding -I (traceroute -I) to make it use ICMP ECHO instead of UDP for probes.
If you want better statistics about packet loss, you can also use a tool called mtr which combines traceroute and ping. If you’re on Ubuntu, you can install this with sudo apt-get install mtr.