AI Integration Engineer
Pratham Adhikari
LLM & Backend Systems / RAG / Real-time Voice Agents
Tandi, Chitwan, Nepal
Profile
who i amBackend and applied-AI engineer with 3 years of production experience owning the AI integration layer end to end: LLM API integration, RAG pipelines, agentic orchestration, output guardrails, and production observability. Shipped real-time voice-to-voice agents at sub-3s end-to-end latency and production RAG systems over regulatory and web-crawled corpora to live users. Deep in Python, FastAPI and containerised microservices; hands-on with pgvector and Solr vector stores, LangChain/LlamaIndex/Pipecat orchestration, Promptfoo-based prompt evaluation, and Langfuse/Grafana/MLflow monitoring of latency, token usage and cost-per-query.
Technical Skills
what i build withWork Experience
3 years shippingAI Engineer · Kathmandu, Nepal
- Owned the AI integration layer across the product, from spec to CI/CD to production observability, integrating Google Gemini APIs and self-hosted LLM, ASR and TTS services into customer-facing backend systems.
- Designed and shipped end-to-end RAG pipelines over heterogeneous corpora — US tax and regulatory documents, and crawled web content — handling ingestion, semantic and recursive chunking, embedding generation, and vector store selection across pgvector and Solr; benchmarked chunking strategies and tuned retrieval parameters against evaluation sets to improve answer relevance.
- Currently architecting the retrieval layer to scale from thousands to millions of documents, evaluating index partitioning, embedding throughput and cost-per-query tradeoffs across vector store options.
- Built agentic and conversational workflows using LangChain, LlamaIndex and Pipecat, including tool-call design, context aggregators and custom orchestration for multi-step flows.
- Established the team's prompt engineering practice: system prompts, few-shot and chain-of-thought patterns, structured output schemas, config-driven prompt versioning, and regression evaluation suites in Promptfoo to catch quality drops before release.
- Implemented user-level style and preference personalisation: typed preference schemas persisted in the database and injected into system prompt construction at request time, giving consistent per-user tone and response style across sessions without retraining.
- Implemented LLM output guardrails in Pipecat — schema-based validation, intent-based routing, branched content filtering and fallback logic — applied to both chat and real-time voice-to-voice agents to prevent unsafe or malformed responses reaching users.
- Instrumented production AI observability with Langfuse, Grafana, MLflow and Dozzle, tracking latency, token usage, cost-per-query and output quality drift to drive optimisation decisions.
- Architected real-time voice-to-voice (V2V) agents orchestrating ASR, LLM and TTS microservices over WebSockets, achieving sub-3-second end-to-end response latency for English and sub-4-second for multilingual conversation through streaming TTS delivery, jitter buffering and pipeline-stage parallelisation.
- Engineered backend microservices in Python/FastAPI with ZeroMQ message-queue orchestration and Docker deployment; implemented authentication and authorisation securing internal APIs.
- Built telephony call infrastructure on Asterisk PBX — multi-language IVR flows and automated inbound/outbound AI-driven call campaigns; load-tested to a fully saturated 16-channel SIP trunk with concurrent live AI agents holding latency targets.
- Fine-tuned Whisper for Nepali and code-mixed ASR, reaching 33% WER on a low-resource language and delivering the organisation's best-performing model; owned the full data pipeline (collection, annotation, filtering, preprocessing). Also fine-tuned TTS and translation models (Chatterbox, Parler-TTS).
- Deployed and maintained on-premise model serving and inference infrastructure alongside Google Cloud services, ensuring data privacy and low-latency performance.
- Led the speech research and backend teams — setting research direction, running code reviews and coordinating delivery across both groups; collaborated with frontend engineers to surface AI capabilities in product UIs.
Projects
side workSemantic search & RAG over exam archives
- Full-stack web app for searching past IOE engineering exam questions, combining lexical and semantic retrieval.
- Built the ingestion pipeline with YOLO for question segmentation and PyTesseract for OCR; indexed content in Elasticsearch with sentence-transformer embeddings for semantic search.
- Extended into a RAG system answering questions from solution documents and textbooks.
Generator / verifier LLM pipeline
- Orchestrated a DeepSeek LLM as generator with a BERT verifier validating generated solutions — an early guardrail/validation pattern for LLM output correctness.
- Applied chain-of-thought prompting and joint answer optimisation to solve complex mathematical problems.
- Trained generator and verifier on the AIMO and Super Mario datasets respectively.
Real-time hand tracking to hardware
- Real-time hand-tracking pipeline using MediaPipe landmarks streamed to a robotic hand for live mimicking.
- Winner, Project Demonstration at Mechtrix 2079 ↗; second place at DELTA 3.0. View demo ↗
Certificates
Education & Fellowships
B.E. Electronics, Information & Communication
Pashchimanchal Campus, IOE, Tribhuvan University ↗ · Pokhara, Nepal
- Ranked 1520 of 15,000 candidates in the 2020 national entrance examination.
- Executive member of the Robotics Club, leading projects on national and international platforms.
Fellow · Kathmandu, Nepal
- Selected among 100+ students for a six-month AI/ML microdegree fellowship, concluding with a capstone project.
Techparva 3.0 — Invited Speaker
Workshop on Datathon · WRC, Nepal
- Delivered sessions introducing machine learning, exploratory data analysis and core ML algorithms.
Research
- Adhikari, P., Bhandari, P., Shrestha, B. & Poudel, A. (2025, January). PBR Map Generation for Realistic 3D Modeling from Images ↗ — generation of textures from real-world objects applicable to any 3D mesh. (Oral presentation)
- Oli, A., Sharma, S., Adhikari, P., Neupane, D., Khanal, S., & Labh, S. K. (2024). Comparative Study on Efficiency Analysis of Fixed and Dual-Axis Solar Tracking System. Journal of Engineering and Sciences, 3(1), 81–86.