Essays & notes
Writing.
Essays on technical leadership, choosing the right problem, and how Hemendra Tripathi thinks about engineering.
Index
- 01Latest
Self-Hosted LLM Inference Crossed a New Threshold, but Tokens/sec Is Still the Wrong Optimization
Production agentic traffic and data-sovereignty constraints are pushing teams past the economic crossover for self-hosted inference. The useful metric is no longer peak tokens/sec.
self-hosted LLMtokens-per-wattagentic workflows
Read essay → - 02
Why the Temporal 2026 Report Crowns the Agentic Engineer (And Spells Doom for Boilerplate Coders)
With 41.1% of AI agents failing in production daily, the highest-leverage career move of 2026 is moving up the stack to own the orchestration layer.
temporal-workflowagentic-workflowssoftware-engineering-careers
- 03
Surviving Gartner’s ‘AI Inference Paradox’: How to Build High-Margin Voice AI Products
As AI token costs drop, startups face a hidden trap: products whose margins shrink as usage scales. Here is how to architect a voice AI stack that grows revenue faster than its infrastructure costs.
voice-aimodel-routingusage-based-billing
- 04
The GPT-6 Astra Tax: Why Lazy Engineering Will Price Out Your Productivity Gains
GPT-6 Astra delivers a massive leap in capability, but treating it as your default model will quickly burn through your engineering budget. The teams that survive the frontier model era are treating model selection as an economic architecture problem.
gpt-6-astrallm-inference-costsmodel-routing
- 05
Redesigning Tech Hiring for Judgment in the Age of AI Sourcing
As AI tools saturate the recruiting funnel with optimized resumes, engineering leaders must shift their focus from automated screening to rigorous, human-led technical calibration.
Technical HiringEngineering LeadershipAI Sourcing
- 06
GPT-6 Astra: Shifting From “Which Model Gives the Best Answer” to Workflow Ownership
Rather than treating frontier models like GPT-6 Astra as a catch-all, teams must adopt a three-tier routing architecture that balances frontier intelligence, utility models, and deterministic code to manage latency, reliability, and cost.
LLM RoutingGPT-6 AstraAI Architecture
- 07
Why Code Writers Are Struggling and How AI Orchestrators Are
If you measure your engineering worth by lines of code, you are competing against a machine that writes syntax infinitely faster. Discover why the future belongs to system owners and AI orchestrators.
AI OrchestrationGenerative AISoftware Engineering
- 08
How We Cut LLM Inference Cost ~20% on a Voice Platform That Scaled Past 1,500 Paying Customers
When every turn of a voice agent is a race against hang-up, one large model destroys margin. Here is the architecture, routing, caching, and billing system that cut LLM spend ~20% while scaling past 1,500 paying customers and nearly eliminating billing disputes.
voice-aillm-cost-optimizationmodel-routing
- 09
Why Some People Are Getting Paid More for the Same Job Title in 2026
Same job title. Different paychecks. In 2026 the gap is no longer just experience or hard work. It is how much leverage people get from cheap, capable AI models. Here is what is actually driving the difference.
Career GrowthAIFuture of Work
- 10
Solving the Right Problem Matters More Than Solving It Perfectly
One of the biggest shifts in my career has been learning that good engineering isn't just about building things well. It's about making sure you're building the right thing in the first place.
EngineeringProduct ThinkingCareer Growth









