AKS
Abhishek*

Senior Software Developer building applied AI systems at scale — multi-agent orchestration, real-time WebRTC platforms, and GPU-backed inference infra powering hundreds of concurrent AI sessions.

Applied AI · Platform Engineering

IamAbhishekKumarSingh,asystems-firstengineer.IshipproductionAI:multi-agentorchestration,real-timeinfra,andGPUinferenceatscale.

OverthelasttwoyearsatRecruit41,Ihaveledthemigrationofareal-timeAIinterviewplatformtoaself-managedLiveKitandGoogleADKstack,scalingto350+concurrentsessionswhilecuttinginfrastructurecostsbyroughly75%andimprovingspeech-processingefficiencybynearly30xonGPU-backedKubernetes.

Experience

Two years shipping production AI.

Recruit41

Bengaluru, India

May 2026 – Present

Senior Software Developer

  • Driving Harness Engineering initiatives through AI-assisted development workflows, evaluation mechanisms, and guardrails to improve code generation accuracy and developer productivity.
  • Designed a hierarchical multi-agent orchestration framework using Google ADK, introducing supervisor agents for tool governance, interview flow control, and candidate evaluation.
  • Led the evolution of a real-time AI interview platform from a custom WebSocket architecture to a WebRTC stack powered by LiveKit and Google ADK.
  • Built an on-the-fly AI evaluation framework that continuously monitored agent behavior, tool usage, and interview progression in production.
  • Architected and operated GPU-backed AI infrastructure on GKE Autopilot for speech and inference workloads.

Recruit41

Bengaluru, India

Sep 2024 – May 2026

Software Developer

  • Spearheaded migration from Daily.co to a self-managed LiveKit deployment supporting 350+ concurrent sessions while reducing infrastructure costs by ~75%.
  • Achieved ~30x speech-processing cost efficiency by replacing external speech APIs with self-hosted inference on GPU-backed Kubernetes workloads.
  • Designed autoscaling infrastructure using Kubernetes HPA and KEDA to dynamically scale communication, speech, and AI workloads during burst traffic.
  • Implemented prompt caching and inference optimization across hosted LLM providers to reduce latency and operating expenses.
  • Conducted load testing and performance tuning of distributed AI systems, resolving scaling bottlenecks in high-concurrency workloads.

Think41

Bengaluru, India

Jul 2024 – Aug 2024

Full Stack Developer Intern

  • Built React and Django applications for an AI-driven interview scheduling platform and recruiter workflow automation.
  • Developed LLM-powered resume screening workflows and implemented automated backend testing using Pytest.
ProductionAIsystemsforteamsthatship.Engineeredforreliability.Poweredbyscale.

Your engineering edge.

Multi-Agent Orchestration. (01)

  • Google ADK supervisor agents for tool governance and flow control.
  • Hierarchical evaluation of candidate and agent behavior in real time.
  • Guardrails for tool usage, safety, and interview progression.
  • Battle-tested across production AI interview workflows.

Real-Time Infra. (02)

  • LiveKit + WebRTC self-managed, 350+ concurrent sessions.
  • GPU inference on GKE Autopilot with HPA and KEDA autoscaling.
  • ~75% cost cut vs third-party providers, ~30x speech efficiency.

AI Evaluation Layer. (03)

  • On-the-fly monitoring of agent behavior and tool usage.
  • Prompt caching and inference optimization across LLM providers.
  • Load-tested pipelines tuned for high-concurrency AI workloads.

Stack

Languages

PythonTypeScriptJavaScript

AI Systems

Google ADKMulti-Agent SystemsAgent OrchestrationAI EvaluationConversational AIPrompt EngineeringPrompt Caching

Infrastructure

KubernetesGKE AutopilotHPAKEDADockerGPU WorkloadsAI InferenceLiveKitWebRTC

Backend

FastAPIDjangoREST APIsMicroservicesWebSockets

Databases

PostgreSQLMongoDBRedis

Open Source

Upstream contributions.

LiveKit Rust SDK

Real-time infrastructure

Diagnosed and resolved an authentication-token propagation issue affecting fallback routing under elevated traffic conditions, and contributed the fix upstream — improving reliability for production LiveKit deployments.

350+

Concurrent AI sessions

~75%

Infra cost reduction

~30x

Speech-processing efficiency

Contact

Let's build something that scales.

Bengaluru, India

© 2026 Abhishek Kumar Singh