Skip to content
Lokesh Jaiswal | Insights
  • Articles
  • About
  • RSS

Pillar

Scalable AI architecture

The infrastructure underneath AI features: retrieval, vector stores, caching, sharding and latency budgets. Each piece works through the arithmetic and the failure modes before the tooling.

  • RAG evaluation: fix retrieval before you blame the model

    Most RAG failures are retrieval failures. A practical playbook for retrieval evaluation: golden sets, recall@k, chunking tests and drift monitoring.

    Scalable AI architecture · 1 October 2026 · 7 min read

  • Embedding model migration: re-indexing 10M+ vectors without downtime

    A better embedding model is only an upgrade if you can install it. The sizing, dual-write and cutover plan for re-indexing 10M+ vectors with no downtime.

    Scalable AI architecture · 1 October 2026 · 9 min read

  • Redis sharding strategy: when a single node stops being the answer

    Most Redis sharding goes wrong before the first shard exists. Diagnose the real ceiling, design keys around hash slots, and reshard without a p99 spike.

    Scalable AI architecture · 1 October 2026 · 10 min read

  • All articles
  • Agentic AI in production
  • Scalable AI architecture
  • EdTech and assessment infrastructure
  • Flutter and mobile at scale
  • Founder and operator lessons

Production notes on AI systems, EdTech and mobile.

  • Agentic AI in production
  • Scalable AI architecture
  • EdTech and assessment infrastructure
  • Flutter and mobile at scale
  • Founder and operator lessons
  • About
  • Privacy
  • RSS feed

© 2026 Lokesh Jaiswal