Technical Lead · AI Engineer · AI Voice EngineerOpen · UTC+05:30

HemendraTripathi

ARCHITECTURE BILLING TEAMS REVENUE

Hemendra Tripathi is a technical lead and AI engineer in Udaipur. He ships as an AI Voice Engineer, plus broader AI product work. He takes products from first commit to paying customers: architecture, billing, teams, and the revenue they produce.

Udaipur, IN "भारत" · US / EU teams · Callin.io Tech Lead

Paying customers
1,500+
LLM inference cost
20%
Voice TTFT (typical)
~420ms
Billing disputes since launch
~0
(01)

Voice AI Case Study

Featured voice AI case study

How we made voice agents cheap enough to scale, and fast enough that humans stayed on the line.

As Technical Lead I owned the architecture, vendor spend, and billing system for an AI voice platform that grew past 1,500 paying customers, including medical and real-estate enterprise accounts, while cutting LLM cost 20% and keeping billing disputes near zero.

The problem
  • Every turn of a voice agent is a race: if the model thinks too long, the caller hangs up.
  • Using one large model for every utterance burned margin on greetings and confirmations.
  • Minute-based billing with rollovers was creating support tickets, and eroding trust with the accounts that mattered most.
Architecture
Production voice pathCallin.io · live turn
  1. 01
    Caller
    PSTN / WebRTC
  2. 02
    Telephony
    Dual-carrier + SIP
  3. 03
    Callin orchestrator
    Cache · parallel prompts · turn clock
  4. 04
    Runtime
    Ultra-low latency · Premium Voice · Custom Stack
  5. 05
    Voice out
    Callin TTS path

We own the path end-to-end. Spend routing and billing sit on the same clock as latency.

What I built

Complexity-aware model routing

Classify each turn (greeting, FAQ, scheduling, objection) and route to the smallest model that can finish the job. Large models only for hard reasoning.

Semantic cache + concurrent prompts

Cache high-frequency intents; fire retrieval and response scaffolds in parallel so time-to-first-token drops before the caller notices silence.

Dual-carrier telephony

Twilio + Telnyx with SIP fallback. Fail over without dropping the call; keep audio streaming over WebSockets under load.

Billing as a product surface

Minute tracking with rollover ledgers precise enough that disputes became rare. Stripe subscriptions wired to actual usage, not estimates.

ReactNode.jsSupabaseTwilioTelnyxStripeElevenLabsCartesiaRedis

Hemendra is the rare engineer who can rewrite the voice pipeline before lunch and close an enterprise prospect after dinner. He treats infrastructure cost like product debt, and it shows in the margins.

(02)

More AI Product Work

Callin.io is the deep dive above. These are related suite products and earlier shipped work. Shorter notes, same bar for outcomes.

01

CondoMail

Production
Product architecture: multi-provider email sync

AI agents that sort, draft, and auto-reply across providers. Live with early adopters, including high-volume Amazon sellers running inbox workflows on it.

React / Node.js / Supabase / Stripe / Firebase
02

Realead

Beta
Full-stack · mobile · AI calling flows

Connects lead sources, builds a business profile, and places AI qualification + follow-up calls for property leads. Final beta ahead of release, and the conversational model behind the agent demo.

React Native / NestJS / Supabase / StripeSee the demo →
03

Sunria & FinTech Accounts

Shipped
Freelance, end-to-end delivery

Pan-India farm management with field-to-warehouse sync, plus a financial dashboard with automated reconciliation: 30% fewer accounting errors, 15+ staff-hours saved weekly.

Laravel / Flutter / MERN / CI-CD
(03)

Experience

OCT 2024 to PRESENT

Technical Lead / AI Engineer / Full-StackAppspundit Infotech · Callin.io

  • Technical Lead: architecture, vendor and infrastructure spend, hiring, product roadmaps, reporting directly to the founder.
  • Scaled the platform to 1,500+ paying customers; expanded into CondoMail and Realead on a shared multi-LLM architecture.
  • Cut LLM inference costs 20%; personally converted 30+ prospects into long-term paying accounts across healthcare and real estate.
2022 to 2024

Freelance Full-Stack DeveloperRemote · fintech, retail, logistics

  • Delivered 10+ end-to-end applications across MERN, Django, Flask, and Laravel.
  • Built a fintech accounts system for an MCA-registered firm, with automated reconciliation saving 15+ staff-hours per week.
  • Improved API performance 25% and implemented zero-downtime CI/CD pipelines.
2021 to 2023

Technical InstructorAimers Institute & VT College

  • Mentored 150+ students in Python, Django, MERN, and Flutter through project-based learning.
  • Redesigned the curriculum to industry needs. Student placements rose 42%.
  • Supervised 30+ capstone projects: version control, API design, UI craft.
BCA, Mohanlal Sukhadia University (2022)English: fluent · Hindi: native
(04)

The Demo Is the Résumé

I build AI voice agents that make real phone calls for a living. This one runs on the same conversational patterns as production voice agents, except its lead-qualification target is you.

Prefer the short path? Read the case study. Prefer proof? Answer the call. The lead file builds as you ask.

  • Same patterns as production qualification agents
  • Finish the call. Get the summary. Email if it lands.
(05)

How I Think About Voice AI

01

Latency is the product

In voice AI, silence is a bug. Every architectural choice (caching, routing, carrier failover) exists to keep the human from hanging up.

02

Pay for intelligence only when you need it

A confirmation doesn't deserve a frontier model. Route by complexity. Your CFO will notice. So will your p95.

03

Billing that doesn't create tickets

If customers argue about invoices, the product is unfinished. Usage ledgers should be boringly correct.

04

Own the stack's P&L

Architecture without vendor spend ownership is theater. I hire, I ship, and I know what the infra bill was last Tuesday.

(06)

Signal

We evaluated three voice vendors. Callin's agents were the only ones our clinic staff didn't hang up on, and the only ones whose invoices we didn't audit line by line.
He hired half my eng team, set the roadmap, and still jumped on customer calls. That's not a contractor. That's an owner.
Curriculum he redesigned moved our placement rate up over 40%. Students left knowing how to ship, not just pass exams.
(07)

AI & Engineering Capabilities

AI & Voice Systems

  • Multi-LLM orchestration: routing by complexity and cost
  • RAG pipelines: Pinecone, Supabase Vector
  • Voice cloning: ElevenLabs, Cartesia
  • Real-time telephony: Twilio, Telnyx, SIP

Product & Revenue

  • Usage-based billing architecture
  • Stripe subscriptions & invoicing
  • Pricing design & infra cost optimization
  • Client acquisition & retention

Full-Stack Engineering

  • React, Next.js, React Native
  • Node.js, Express, NestJS, Laravel, Python
  • PostgreSQL (Supabase), Redis
  • Event-driven systems, REST, microservices

Cloud & Leadership

  • Docker, AWS, CI/CD pipelines
  • System design & architecture decisions
  • Team leadership: hiring and mentorship
  • US/EU stakeholder coordination
(08)

Off the Record

Hemendra Tripathi, technical lead and AI engineer

Based in Udaipur, shipping for the US and Europe. I got here by teaching 150+ students to code, freelancing across four frameworks, and rebuilding a voice-AI platform until 1,500 companies paid for it. I like systems that are boring, fast, and profitable, and teams that own what they build.
I’m Batman.

(09)

FAQ

Who is Hemendra?

Hemendra is Hemendra Tripathi, a technical lead and AI engineer in Udaipur, Rajasthan. He leads Callin.io at Appspundit Infotech as an AI Voice Engineer: multi-LLM voice agents used by 1,500+ paying customers. His official site is https://me.readwith.io.

Who is Hemendra Tripathi?

Hemendra Tripathi is a technical lead and AI engineer based in Udaipur, shipping for US and EU teams. As an AI Voice Engineer and Technical Lead on Callin.io he scaled multi-LLM voice agents past 1,500 paying customers, cut inference cost about 20%, and owns architecture, billing, and hiring.

What is an AI Voice Engineer?

An AI Voice Engineer designs and ships real-time voice agents: telephony, model routing, latency, and billing. Hemendra Tripathi is an AI Voice Engineer and technical lead in Udaipur. At Callin.io he scaled multi-LLM voice agents past 1,500 paying customers and cut inference cost about 20%.

Which Hemendra Tripathi is the Callin.io technical lead?

Hemendra Tripathi of Udaipur, Rajasthan is the technical lead and AI engineer at Appspundit Infotech behind Callin.io. His official site is https://me.readwith.io. He is not the Newstrack journalist or other professionals who share the same name.

How do I hire an AI Voice Engineer like Hemendra?

Book a 20-minute call at cal.com/hemendratripathi/hiring-freelance. Timezone is converted for you. Or email hemendratripathi880@gmail.com. He is open to technical-lead, AI engineer, and senior full-stack roles, plus select freelance. Expect a reply within 24 hours, plus the Callin.io case study.

What is complexity-aware multi-LLM orchestration?

It classifies each voice-agent turn (greeting, FAQ, scheduling, objection) and routes to the smallest model that can finish the job. Large models handle hard reasoning only. Paired with semantic cache and concurrent prompts, this cuts LLM spend while keeping time-to-first-token low enough callers stay on the line.

What voice AI stack does Hemendra ship with?

Production systems use React and Node.js with Twilio, Telnyx, and SIP failover, plus ElevenLabs or Cartesia for voice, Supabase or Redis for state and cache, and Stripe for usage-based billing. RAG paths use Pinecone or Supabase Vector when retrieval is required.

What results did the Callin.io voice AI platform achieve?

Under Hemendra’s technical leadership the platform grew past 1,500 paying customers, including healthcare and real-estate accounts. Model routing cut LLM inference cost about 20%, typical voice TTFT landed near 420ms, and precise minute ledgers kept billing disputes near zero after launch.

Is Hemendra available for freelance or full-time AI engineering roles?

Yes. He is open to technical-lead, AI engineer, and senior full-stack roles, and select freelance focused on voice AI, multi-LLM products, or usage-based billing. Book 20 minutes at cal.com/hemendratripathi/hiring-freelance or email; he typically replies within 24 hours and can share the Callin.io case study and résumé.

(10) Contact

HIRE ME.

Looking for a technical lead and AI engineer who has already shipped voice AI into revenue, and can stretch into other AI product work. Open to technical-lead / AI engineer / senior full-stack roles and select freelance. Pick a 20-minute slot (timezone is converted for you), or email if you prefer. Replies within 24 hours.

US / EU-friendly hours · timezone handled