Senior SDE · Amazon Alexa AI

Harshit Joshi

I build scalable, low-latency distributed systems and the GenAI/LLM orchestration that powers conversational AI for millions of people.

View Technical Portfolio

About

I am a senior engineer focused on the systems that keep voice-first AI fast and reliable. My work sits where high-throughput distributed infrastructure meets large language model orchestration: routing, batching, and serving inference so that response times stay in the tens of milliseconds even under production load.

I care about the unglamorous parts that make scale possible - backpressure, graceful degradation, observability, and the data paths that let a model answer correctly the first time.


Experience

Six years at Amazon, spanning three teams across the Alexa and retail organizations.

  • Alexa AIAmazon
    Senior Software Development Engineer

    GenAI/LLM orchestration and low-latency inference at conversational scale.

  • Alexa NotificationsAmazon
    Software Development Engineer

    High-throughput delivery pipelines fanning out to hundreds of millions of devices.

  • Search & CatalogAmazon
    Software Development Engineer

    Distributed indexing and retrieval systems across a planet-scale product catalog.


Focus


Writing

Deep dive · Local inference

Model Optimization Parameters using llama.cpp

Running a capable model on your own hardware is mostly a game of tradeoffs: memory against quality, throughput against latency, CPU against GPU. Here is how I tune llama.cpp to get the most out of a single machine.


Ask

A quick way to get the gist before you reach out. Ask the assistant about my work, stack, or availability - answers are written in my voice, kept short on purpose.

recruiter-assistant
harshit

Hi - I'm a stand-in for Harshit. Ask me about his work, his stack, or whether he's open to new roles. Pick a prompt below or type your own.