MS Mukul Sharma
Resume
Tech Lead · AI & Backend Engineering · Meerut, India

I build production AI systems that talk, read and decide.

5+ years architecting real-time voice AI agents, LLM document intelligence pipelines, vector DB search systems, and vLLM GPU inference microservices for enterprise workflows.

Signal Pipeline v2.4
STATUS: STREAMING
RAW INPUT AUDIO Plivo WebSocket
8kHz PCM Audio Chunk: 20ms
STRUCTURED CRM PAYLOAD Gemini Live + vLLM
{
  "lead": "Enterprise Client",
  "intent": "RFP_Qualified",
  "confidence": 0.98,
  "action": "SYNC_SALESFORCE"
}
vLLM GPU Queue
Vector Index Synced
5+ Yrs
Backend & AI Engineering
1,000+
Docs Processed / Month
95%
Workflows Automated
15+
Integrations Shipped
PROVED IN PRODUCTION

Selected Case Studies

Detailed breakdown of architectural decisions, constraints, trade-offs, and measurable outcomes.

01 // CASE STUDY Real-Time Voice AI

Real-Time AI Voice Lead-Qualification Bot

Sub-second bidirectional voice pipeline connecting Gemini Live, Plivo telephony, and MiClient CRM.

Architected a real-time voice AI qualification system that ingests inbound phone leads via Plivo, streams bidirectional audio into Gemini Live, and streams structured CRM updates.

Google Gemini Live Plivo Telephony Python / FastAPI WebSockets CRM Integration
Read architectural case study
02 // CASE STUDY 95% Manual Workflow Automated

AI RFP Processing & Bidding Workflow

Automated PDF & scanned document requirement extraction, product-customer matching, deal prediction, and confidence scoring.

Built an intelligent Request for Proposal (RFP) engine that parses complex multi-page PDF/scanned documents, extracts granular line-item requirements, and predicts deal win probability.

LLM Document Intelligence Requirement Extraction Deal Prediction Python / FastAPI Node.js
Read architectural case study
03 // CASE STUDY 1,000+ Docs / Month

Document Intelligence for CRM Platform

High-throughput document ingestion engine processing 1,000+ documents/month with 15+ third-party CRM integrations.

Architected a scalable document intelligence pipeline processing 1,000+ invoices, purchase orders, and customer contracts monthly while integrating seamlessly with Gmail, Salesforce, QuickBooks, and WhatsApp.

Document Processing Integration Hub Salesforce QuickBooks Stripe Razorpay
Read architectural case study
04 // CASE STUDY vLLM & RAG Infrastructure

Vector Search & GPU Inference Infrastructure

High-performance vLLM GPU cluster, vector embedding pipelines, AWS-to-GCP modernization, and hybrid FastAPI + Node.js services.

Designed and modernized the AI core backend, transitioning workloads from AWS to GCP, deploying self-hosted vLLM GPU inference instances, and building scalable RAG vector search pipelines.

vLLM GPU Serving Vector Database RAG Pipeline AWS to GCP Migration FastAPI / Node.js
Read architectural case study
SYSTEM ARCHITECTURE

How I Build Production AI

End-to-end signal pipeline engineering: from edge client connections to GPU inference clusters and CRM state mutation.

SIGNAL PIPELINE DATAFLOW Client ──> API ──> Queue ──> LLM ──> Vector DB ──> CRM
STAGE 01 Client Ingestion Web, Voice, Webhooks
STAGE 02 API Gateway FastAPI / Node.js
STAGE 03 Worker Queue Async Pipeline
STAGE 04 (GPU) vLLM & Gemini Model Inference
STAGE 05 Vector DB Embeddings & RAG
STAGE 06 MiClient CRM Structured State
01 // PRINCIPLE

Reliability First

AI components are probabilistic; backend systems must be deterministic. Pydantic schemas, fallbacks, and retry queues guarantee system integrity.

02 // PRINCIPLE

Granular Observability

Tracing every token, audio chunk, and API callback with per-request correlation IDs to pinpoint latency anomalies and model degradation.

03 // PRINCIPLE

Cost & Latency Budgeting

Self-hosting vLLM GPU inference instances and deploying hybrid semantic cache layers to keep p95 latency low and token costs predictable.

04 // PRINCIPLE

Continuous Evaluation

Automated regression testing on extraction pipelines using synthetic dataset benchmarks before promoting prompt changes or model weights.

ECOSYSTEM & APIS

15+ Enterprise Integrations & AI Stack

Production-tested webhooks, REST APIs, OAuth workflows, and streaming WebSocket pipelines built for MiClient CRM/CPQ platform.

Google Gemini Live
Voice & Real-Time AI
Plivo Telephony
Voice WebSockets
vLLM GPU Serving
Inference Engine
Vector DB / RAG
Semantic Search
Python / FastAPI
AI Services
Node.js
Backend Microservices
Salesforce CRM
CRM Integration
QuickBooks
Financial Sync
Stripe
Payments
Razorpay
Payments
Meta / WhatsApp
Messaging
Microsoft Graph / Teams
Enterprise Communication
Gmail API
Email Processing
SendGrid
Transactional Mail
JustDial & IndiaMART
B2B Lead Channels
CAREER PATH

Engineering Experience

5+ years leading backend & AI architecture across high-growth startups and tech enterprises.

Tech Lead, Engineering @ MiClient Pvt. Ltd.

Jul 2025 – Present
Noida / Meerut, India
  • Leading core backend & AI engineering teams across voice AI, document intelligence, and cloud microservices.
  • Architected sub-second real-time voice lead qualification bot using Google Gemini Live + Plivo WebSockets.
  • Spearheaded AWS-to-GCP infrastructure modernization, cutting GPU inference overhead and scaling vLLM clusters.

Software Development Engineer @ MiClient Pvt. Ltd.

Dec 2022 – Jun 2025
Noida / Meerut, India
  • Engineered automated AI RFP bidding workflow processing multi-page PDFs with requirement extraction & deal scoring.
  • Designed document intelligence engine processing 1,000+ documents/month with 95% workflow automation.
  • Built 15+ third-party integrations (Salesforce, QuickBooks, Stripe, Razorpay, Meta/WhatsApp, Teams, SendGrid).

Software Development Engineer (SDE) @ Amazon

Jan 2022 – Sep 2022
Chennai, India
  • Engineered core system features for Fire OS Settings, Keyboard, and Parental Control services on Fire TV devices.
  • Optimized Android/Fire OS system frameworks ensuring high-reliability input handling and low memory footprint.

Software Engineer @ Nagarro

Jun 2021 – Jan 2022
Gurugram, India
  • Developed enterprise web services and backend APIs adhering to microservices design patterns.

B.Tech in Computer Science & Engineering @ Indian Institute of Information Technology (IIIT) Bhagalpur

2017 – 2021
Bhagalpur, India
  • Graduated in CSE. Deep focus on algorithms, system design, operating systems, and distributed computing.
COMPETITIVE & RECOGNITION

Beyond Work & Problem Solving

CodeChef

4★ Rated (Max Rating: 1955)

Competitive programming focused on algorithmic optimization, dynamic programming, and graph algorithms.

Codeforces

Specialist (Max Rating: 1561)

Problem solving under strict time and space complexity constraints.

Smart India Hackathon 2020

2nd Place (Institute Level)

Built scalable tech solution addressing real-world problem statement under 36-hour hackathon environment.

Interactive Signal Sandbox (Demo Stub)
RATE-LIMITED · MOCK PREVIEW

Test raw text chunking and structured JSON extraction. Demonstrates client payload parsing with rate-limiting validation.

{ "intent": "RFP_MODULE_ENQUIRY", "integrations": ["QuickBooks"], "score": 0.96 }
Backend API: FastAPI + vLLM (Mocked client stub) Ready for integration
GET IN TOUCH

Let's build something that works in production.

Whether you're looking for a Tech Lead to architect real-time voice AI, build document intelligence systems, or scale vLLM GPU inference infrastructure, feel free to reach out.