Happy Monday! ☀️

Welcome to the 395 new hungry minds who have joined us since last Monday!

If you aren't subscribed yet, join smart, curious, and hungry folks by subscribing here.

Booking.com runs 70+ Node.js services for server-side rendering, assembling pages from React components owned by different teams. These services had been configured the same way for years—running four to eight workers via pm2—without anyone questioning whether those defaults still made sense. When they finally took a holistic look at the entire stack, they found substantial room to optimize.

The real challenge wasn't obvious from aggregated metrics. Teams were hitting latency targets, resource usage looked reasonable, and nothing triggered alerts. But beneath the surface, years of platform growth and shifting workloads had made the original configuration suboptimal. By testing systematically and measuring deeply, they discovered they could run the same traffic with 30% fewer pods, use 20% less memory per pod, and cut latency at p75/p99 by up to 10%.

The challenge: Node's single-threaded event loop and process-based worker isolation were leaving CPU underutilized and adding unnecessary overhead on every request.

Implementation highlights:

  1. Switched from pm2 to wattpm. Replaced OS processes with worker threads and kernel-level connection distribution via SO_REUSEPORT, eliminating the supervisor hop on each request and reducing memory duplication at startup.

  2. Fixed uneven load distribution. Discovered that kernel connection reuse combined with hash-based worker selection created 20–40% imbalance. Solved it by having nginx load-balance across workers on different ports instead.

  3. Tested rigorously before production. Ran synthetic load tests under three scenarios (normal load, extreme stress, mixed workload) to prove wattpm handled 40–50% more requests at the same latency on realistic traffic patterns.

  4. Split traffic at the routing layer. Deployed two complete identical installations side-by-side in production and routed 50/50 traffic between pm2 and wattpm, keeping pod counts and resources equal to isolate the change.

  5. Made it opt-in via feature flag. Changes flipped at startup, not per-request, so 70+ teams could adopt safely with a documented rollback path if needed.

Results and learnings:

  • 38% cost reduction came from running 30% fewer pods and cutting 20% memory per pod. No rewrites, no new hardware—pure optimization through better use of existing resources.

  • Aggregated metrics hid the problem. Per-worker visibility revealed load imbalance that p50/p75/p90 across a whole pod completely masked, showing why local profiling matters at scale.

  • Production contradicts synthetic tests. Moderate-load latency regressed in prod even though labs showed gains, forcing them to dig deeper and discover the kernel connection-reuse issue their tests never encountered.

The lesson: sometimes the highest ROI optimization isn't a new algorithm or a clever trick—it's removing assumptions that haven't been questioned since they were set. Just be ready to debug when your synthetic benchmarks and production reality decide to play hide-and-seek with each other.

ARTICLE (rust serialization reimagined)
Deser: Rethinking Rust Serialization

ARTICLE (embedding vectors at scale)
Evolving Pinterest's Embedding Retrieval Platform

ARTICLE (2026 llm chaos)
2026 in LLMs (so far)

ARTICLE (ai agents breaking bad)
An Agent Used DNS to Reach an External Chatbot

ARTICLE (data is forever)
Data is the application

ARTICLE (chatgpt api integration)
Using Your ChatGPT Plan in Other Apps and Sites

Want to reach 200,000+ engineers?

Let’s work together! Whether it’s your product, service, or event, we’d love to help you connect with this awesome community.

Brief: OpenAI released GPT-6.1 Sol, matching flagship GPT-6 Astra performance on coding and professional tasks at one-fifth the cost with cached inputs at $0.10 per million tokens.

Brief: OpenAI unveiled a Decisions API enabling applications to route and classify tasks using GPT-6 Luna at DevDay 2026.

Brief: Anthropic's new Claude Sonnet 5.5 runs 30% faster and costs up to 30% less than its predecessor while improving performance across coding and knowledge work tasks.

Brief: Claude solved a nine-loop particle physics calculation in N=4 super Yang-Mills theory, breaking through a computational barrier that researchers expected would require more resources than typically available.

Brief: Google's new Gemini 4 Argon frontier model supports 1M token output and excels in software engineering, enterprise workflows, and cybersecurity defense.

Brief: xAI released Team Bots that learn and coordinate work inside Slack, enabling shared AI teammates with context, plugins, and memory across sales, engineering, and marketing workflows.

This week’s tip:

Query Result Streaming and Cursor-Based Pagination for Memory Efficiency

Using database cursors and streaming result sets prevents loading entire result sets into memory, critical for querying large tables or building responsive paginated APIs without OOM events or request timeouts.

Wen?

  • When exporting millions of records to prevent API server memory exhaustion

  • When pagination UX requires stable cursors that survive concurrent data modifications

  • When building event streaming systems that need constant memory footprint regardless of query result size

People do not decide their futures, they decide their habits and their habits decide their futures.
Gary Keller

That’s it for today! ☀️

Enjoyed this issue? Send it to your friends here to sign up, or share it on Twitter!

If you want to submit a section to the newsletter or tell us what you think about today’s issue, reply to this email or DM me on Twitter! 🐦

Thanks for spending part of your Monday morning with Hungry Minds.
See you in a week — Alex.

Icons by Icons8.

*I may earn a commission if you get a subscription through the links marked with “aff.” (at no extra cost to you).