Pratham Adhikari

AI Integration Engineer

Pratham Adhikari

LLM & Backend Systems / RAG / Real-time Voice Agents

Tandi, Chitwan, Nepal

01

Profile

who i am

Backend and applied-AI engineer with 3 years of production experience owning the AI integration layer end to end: LLM API integration, RAG pipelines, agentic orchestration, output guardrails, and production observability. Shipped real-time voice-to-voice agents at sub-3s end-to-end latency and production RAG systems over regulatory and web-crawled corpora to live users. Deep in Python, FastAPI and containerised microservices; hands-on with pgvector and Solr vector stores, LangChain/LlamaIndex/Pipecat orchestration, Promptfoo-based prompt evaluation, and Langfuse/Grafana/MLflow monitoring of latency, token usage and cost-per-query.

< 3 send-to-end voice-to-voice latency
33 % WERNepali ASR — org's best model
16-chSIP trunk saturated with live AI agents
02

Technical Skills

what i build with
03

Work Experience

3 years shipping

Wiseyak

AI Engineer · Kathmandu, Nepal

  • Owned the AI integration layer across the product, from spec to CI/CD to production observability, integrating Google Gemini APIs and self-hosted LLM, ASR and TTS services into customer-facing backend systems.
  • Designed and shipped end-to-end RAG pipelines over heterogeneous corpora — US tax and regulatory documents, and crawled web content — handling ingestion, semantic and recursive chunking, embedding generation, and vector store selection across pgvector and Solr; benchmarked chunking strategies and tuned retrieval parameters against evaluation sets to improve answer relevance.
  • Currently architecting the retrieval layer to scale from thousands to millions of documents, evaluating index partitioning, embedding throughput and cost-per-query tradeoffs across vector store options.
  • Built agentic and conversational workflows using LangChain, LlamaIndex and Pipecat, including tool-call design, context aggregators and custom orchestration for multi-step flows.
  • Established the team's prompt engineering practice: system prompts, few-shot and chain-of-thought patterns, structured output schemas, config-driven prompt versioning, and regression evaluation suites in Promptfoo to catch quality drops before release.
  • Implemented user-level style and preference personalisation: typed preference schemas persisted in the database and injected into system prompt construction at request time, giving consistent per-user tone and response style across sessions without retraining.
  • Implemented LLM output guardrails in Pipecat — schema-based validation, intent-based routing, branched content filtering and fallback logic — applied to both chat and real-time voice-to-voice agents to prevent unsafe or malformed responses reaching users.
  • Instrumented production AI observability with Langfuse, Grafana, MLflow and Dozzle, tracking latency, token usage, cost-per-query and output quality drift to drive optimisation decisions.
  • Architected real-time voice-to-voice (V2V) agents orchestrating ASR, LLM and TTS microservices over WebSockets, achieving sub-3-second end-to-end response latency for English and sub-4-second for multilingual conversation through streaming TTS delivery, jitter buffering and pipeline-stage parallelisation.
  • Engineered backend microservices in Python/FastAPI with ZeroMQ message-queue orchestration and Docker deployment; implemented authentication and authorisation securing internal APIs.
  • Built telephony call infrastructure on Asterisk PBX — multi-language IVR flows and automated inbound/outbound AI-driven call campaigns; load-tested to a fully saturated 16-channel SIP trunk with concurrent live AI agents holding latency targets.
  • Fine-tuned Whisper for Nepali and code-mixed ASR, reaching 33% WER on a low-resource language and delivering the organisation's best-performing model; owned the full data pipeline (collection, annotation, filtering, preprocessing). Also fine-tuned TTS and translation models (Chatterbox, Parler-TTS).
  • Deployed and maintained on-premise model serving and inference infrastructure alongside Google Cloud services, ensuring data privacy and low-latency performance.
  • Led the speech research and backend teams — setting research direction, running code reviews and coordinating delivery across both groups; collaborated with frontend engineers to surface AI capabilities in product UIs.
04

Projects

side work

PrashnaKhoj

Semantic search & RAG over exam archives

  • Full-stack web app for searching past IOE engineering exam questions, combining lexical and semantic retrieval.
  • Built the ingestion pipeline with YOLO for question segmentation and PyTesseract for OCR; indexed content in Elasticsearch with sentence-transformer embeddings for semantic search.
  • Extended into a RAG system answering questions from solution documents and textbooks.
PythonPyTorchElasticsearch ReactNode.js

Intelli-Math

Generator / verifier LLM pipeline

  • Orchestrated a DeepSeek LLM as generator with a BERT verifier validating generated solutions — an early guardrail/validation pattern for LLM output correctness.
  • Applied chain-of-thought prompting and joint answer optimisation to solve complex mathematical problems.
  • Trained generator and verifier on the AIMO and Super Mario datasets respectively.
DeepSeekBERTChain-of-thought

ML-Based Robotic Arm

Real-time hand tracking to hardware

MediaPipeOpenCVEmbedded
05

Certificates

06

Education & Fellowships

B.E. Electronics, Information & Communication

Pashchimanchal Campus, IOE, Tribhuvan University · Pokhara, Nepal

  • Ranked 1520 of 15,000 candidates in the 2020 national entrance examination.
  • Executive member of the Robotics Club, leading projects on national and international platforms.

Fusemachines AI Fellowship

Fellow · Kathmandu, Nepal

  • Selected among 100+ students for a six-month AI/ML microdegree fellowship, concluding with a capstone project.

Techparva 3.0 — Invited Speaker

Workshop on Datathon · WRC, Nepal

  • Delivered sessions introducing machine learning, exploratory data analysis and core ML algorithms.
07

Research

  • Adhikari, P., Bhandari, P., Shrestha, B. & Poudel, A. (2025, January). PBR Map Generation for Realistic 3D Modeling from Images — generation of textures from real-world objects applicable to any 3D mesh. (Oral presentation)
  • Oli, A., Sharma, S., Adhikari, P., Neupane, D., Khanal, S., & Labh, S. K. (2024). Comparative Study on Efficiency Analysis of Fixed and Dual-Axis Solar Tracking System. Journal of Engineering and Sciences, 3(1), 81–86.
08

Contact & References

Suresh Manandhar, Ph.D. CEO / Chief Scientist / Professor · Wiseyak, Nepal smanandhar@york.ac.uk
Khem Raj Koirala HOD, Dept. of Computer & Electronics Engineering · Pashchimanchal Campus, IOE, Nepal doece@ioepas.edu.np