OPTIMUS PROJECT / architects.lab
System Design Β· Internal Platform

Self-Hosted AI Code Generation Platform
for a 10,000-Person Fintech Org

Replacing expensive third-party AI coding assistants with self-hosted LLMs in a regulated fintech environment β€” built for data sovereignty, cost control, and production reliability, and wired directly into the Jira workflow via an autonomous agent.

Scale: 10,000 devs
Savings: 54–72%
Cloud: AWS
Primary model: DeepSeek Coder 33B
Timeline: 18 months
00

Executive Summary

Recall first β€” before you read on
  • What third-party AI coding tools were being replaced, and why?
  • What three priorities does the solution claim to optimize for?
Executive SummaryOVERVIEW
β–Ά

This document presents a complete system architecture for replacing expensive third-party AI coding assistants with self-hosted Large Language Models in a regulated fintech environment. The solution prioritizes data sovereignty, cost control, and production reliability while serving 10,000 developers across multiple teams.

What we learned

A regulated fintech can replace third-party AI coding tools with a self-hosted stack without giving up developer productivity β€” as long as sovereignty, cost, and reliability are designed in from day one, not bolted on later.

01

The Business Case: Why Self-Hosting Wins

Recall first β€” before you read on
  • Roughly what does the current third-party API spend look like per year for 10,000 devs?
  • What's the ballpark annual cost of the self-hosted alternative, and the resulting savings range?
The Cost Reality CheckFINANCE
β–Ά

Let me paint you a picture of what's actually happening in your organization right now:

cost-comparison.txtascii
Current State (Third-Party APIs):
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  10,000 Developers Γ— ~50 API calls/day Γ— $0.03/call        β”‚
β”‚  = $15,000/day = $450,000/month = $5.4M/year               β”‚
β”‚                                                             β”‚
β”‚  But wait - this is conservative. In reality:               β”‚
β”‚  β€’ Heavy users make 200+ calls/day                         β”‚
β”‚  β€’ Code generation requires longer context windows         β”‚
β”‚  β€’ Premium models cost $0.05-0.10 per call                 β”‚
β”‚  β€’ Real cost: $8-12M/year                                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Self-Hosted Alternative:
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Infrastructure (Reserved Instances, 3-year commitment):    β”‚
β”‚  β€’ 8Γ— p4d.24xlarge (A100-40GB): ~$32/hour each             β”‚
β”‚  β€’ Total compute: ~$256/hour = $186,880/month              β”‚
β”‚  β€’ Storage, networking, management: ~$50,000/month         β”‚
β”‚  β€’ Engineering team (3 FTE): ~$50,000/month                β”‚
β”‚  Total: ~$286,880/month = $3.4M/year                       β”‚
β”‚                                                             β”‚
β”‚  Annual Savings: $4.6-8.6M (54-72% reduction)               β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
CTO'S REAL QUESTIONIs cost the only driver? No. Here's what keeps me awake at night β€”
What we learned

At 10,000-developer scale, per-call API pricing scales linearly with usage while self-hosted infrastructure scales in steps (GPU nodes) β€” so the larger and heavier the usage, the more self-hosting's economics win.

Recall first β€” before you read on
  • Besides cost, what's the other major driver for self-hosting?
  • What does code leaving the VPC expose the bank to that self-hosting avoids?
The Control ImperativeGOVERNANCE
β–Ά
data-sovereignty-matrix.txtascii
Data Sovereignty Matrix:
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    β”‚ Third-Party API β”‚ Self-Hosted         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Code leaves VPC    β”‚ YES ⚠️          β”‚ NO βœ…               β”‚
β”‚ Provider sees code β”‚ YES ⚠️          β”‚ NO βœ…               β”‚
β”‚ Audit trail        β”‚ Limited         β”‚ Complete            β”‚
β”‚ Model customizationβ”‚ Impossible      β”‚ Full control        β”‚
β”‚ Latency SLA        β”‚ Provider's      β”‚ Your infrastructure β”‚
β”‚ Regulatory (SOC2)  β”‚ Complex         β”‚ Direct control      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
THOUGHT-PROVOKING QUESTIONIf your competitor gets acquired by your AI provider, what happens to all the proprietary code patterns that have been sent to their API?
What we learned

Cost savings alone don't justify the migration risk β€” the real unlock is that code never leaves the VPC, giving complete audit trails and regulatory control that a third-party API can never offer.

02

LLM Selection: The Evaluation War Room

Recall first β€” before you read on
  • Which 8 candidate models were evaluated, and across how many dimensions?
  • Which model scored highest on raw code-gen quality β€” and was it the one finally chosen?
The Candidate Pool8 MODELS Β· 6 DIMENSIONS
β–Ά

We evaluated 8 models across 6 dimensions. Here's the brutal truth:

evaluation-matrix.txtscale 1–10
Evaluation Matrix (Scale 1-10, higher = better):
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Model                β”‚ Code Gen β”‚ Latency  β”‚ Memory   β”‚ Fine-tuneβ”‚ License  β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Code Llama 34B       β”‚    8.2   β”‚   7.5    β”‚   8.0    β”‚   9.0    β”‚  Open    β”‚
β”‚ DeepSeek Coder 33B   β”‚    8.5   β”‚   7.8    β”‚   8.2    β”‚   8.5    β”‚  Open    β”‚
β”‚ Mistral 7B           β”‚    6.5   β”‚   9.5    β”‚   9.5    β”‚   8.0    β”‚  Apache  β”‚
β”‚ StarCoder2 15B       β”‚    7.8   β”‚   8.0    β”‚   8.5    β”‚   7.5    β”‚  Open    β”‚
β”‚ Llama 2 70B          β”‚    8.8   β”‚   5.0    β”‚   5.0    β”‚   9.0    β”‚  Custom  β”‚
β”‚ GPT-NeoX 20B         β”‚    6.0   β”‚   7.0    β”‚   7.0    β”‚   7.0    β”‚  Apache  β”‚
β”‚ Phi-2 (fine-tuned)   β”‚    7.2   β”‚   9.0    β”‚   9.8    β”‚   6.0    β”‚  MIT     β”‚
β”‚ WizardCoder 34B      β”‚    8.7   β”‚   7.0    β”‚   7.5    β”‚   8.0    β”‚  Open    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Benchmark scores on paper (Code Gen, Latency, Memory, Fine-tune, License) rarely tell the full story β€” no single model wins on every axis, which is exactly why a multi-model strategy made more sense than picking one 'best' model.

Recall first β€” before you read on
  • What three-model strategy was chosen β€” a primary, a fast path, and a specialized path?
  • Why was Llama 2 70B rejected despite strong benchmark scores?
  • How did the internal 500-PR benchmark differ from public benchmarks?
The Final Decision: A Multi-Model StrategyDECISION
β–Ά
model-decision-tree.txtascii
WHY WE CHOSE WHAT WE CHOSE:

Primary Workhorse: DeepSeek Coder 33B
β”œβ”€β”€ Best code generation quality-to-latency ratio
β”œβ”€β”€ Open license with no commercial restrictions
β”œβ”€β”€ Excellent at understanding financial domain code patterns
└── Run on A100-40GB with comfortable headroom

Fallback/Fast Path: Mistral 7B (Fine-tuned)
β”œβ”€β”€ Lightning fast for simple completions
β”œβ”€β”€ Can run on T4 GPUs at fraction of cost
β”œβ”€β”€ Perfect for IDE autocomplete scenarios
└── 95% of simple queries don't need 33B parameters

Specialized Path: WizardCoder 34B
β”œβ”€β”€ Used for complex refactoring tasks
β”œβ”€β”€ Better at test generation
└── Only invoked by explicit user request

WHY OTHERS WERE REJECTED:

Llama 2 70B: Too large, too slow
β”œβ”€β”€ 2.5s inference time on A100 (unacceptable for IDE use)
β”œβ”€β”€ Requires 2Γ— A100s for comfortable deployment
└── Meta's custom license created legal review delays

GPT-NeoX: Outdated architecture
β”œβ”€β”€ RoPE embeddings missing, context handling poor
└── Community support declining

StarCoder2: Good but not great
β”œβ”€β”€ Filled-attention bugs in initial release
└── Finetuning documentation sparse

Phi-2: Too small for complex tasks
β”œβ”€β”€ Failed on multi-file refactoring
└── Cannot handle our 16K context requirements
Internal Benchmark β€” Did we test on OUR code? β–Ά

Yes. We built an internal benchmark of 500 real PRs from the last year:

internal-benchmark.txtPass@1
Internal Benchmark Results (Pass@1):
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Task Category          β”‚ DS-Coder β”‚ WizCoder β”‚ Mistral β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ CRUD API Generation    β”‚   92%    β”‚   89%    β”‚   78%   β”‚
β”‚ SQL Query Building     β”‚   88%    β”‚   85%    β”‚   82%   β”‚
β”‚ Unit Test Writing      β”‚   76%    β”‚   82%    β”‚   65%   β”‚
β”‚ Refactoring (Multi-file)β”‚   71%    β”‚   74%    β”‚   42%   β”‚
β”‚ Regex/Pattern Match    β”‚   94%    β”‚   91%    β”‚   88%   β”‚
β”‚ Security Review        β”‚   68%    β”‚   65%    β”‚   55%   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

The best model for a task isn't always the biggest one: routing 95% of simple completions to a small fine-tuned model while reserving the 33B model for real generation work is what actually made the economics and latency work.

03

AWS Infrastructure: The Physical Reality

Recall first β€” before you read on
  • Which AWS network layers sit in front of the wrapper service β€” public vs private vs GPU compute subnets?
  • Where do SageMaker endpoints for DeepSeek and Mistral sit in the VPC?
Topology OverviewVPC / NETWORK
β–Ά
vpc-topology.txtascii
                    INTERNET
                       β”‚
                       β–Ό
              [AWS WAF + Shield]
                       β”‚
                       β–Ό
              [Route 53 DNS]──────[CloudFront CDN]
                       β”‚                  β”‚
                       β–Ό                  β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚     VPC (10.0.0.0/16)             β”‚
              β”‚                                    β”‚
              β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
              β”‚  β”‚   Public Subnets (DMZ)        β”‚ β”‚
              β”‚  β”‚                               β”‚ β”‚
              β”‚  β”‚  [Application Load Balancer]  β”‚ β”‚
              β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚ β”‚
              β”‚  β”‚  β”‚  ALB-1a  β”‚ β”‚  ALB-1b  β”‚   β”‚ β”‚
              β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚ β”‚
              β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
              β”‚                 β”‚                  β”‚
              β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
              β”‚  β”‚   Private App Subnets        β”‚ β”‚
              β”‚  β”‚                              β”‚ β”‚
              β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ β”‚
              β”‚  β”‚  β”‚  IDFC-Coder Wrapper    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  (ECS Fargate)         β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”‚Task-1β”‚ β”‚Task-2β”‚    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚ β”‚
              β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚ β”‚
              β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
              β”‚              β”‚                     β”‚
              β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
              β”‚  β”‚   GPU Compute Subnets        β”‚ β”‚
              β”‚  β”‚                              β”‚ β”‚
              β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ β”‚
              β”‚  β”‚  β”‚  SageMaker Endpoints   β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚                        β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  Endpoint-1 (DS-33B):  β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”‚ml.p4dβ”‚ β”‚ml.p4dβ”‚    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”‚ .24xlβ”‚ β”‚ .24xlβ”‚    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚                        β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  Endpoint-2 (Mistral): β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”‚ml.g5 β”‚ β”‚ml.g5 β”‚    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β”‚ .12xlβ”‚ β”‚ .12xlβ”‚    β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜    β”‚  β”‚ β”‚
              β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚ β”‚
              β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
              β”‚                                    β”‚
              β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
              β”‚  β”‚   Data/State Subnets         β”‚ β”‚
              β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ β”‚
              β”‚  β”‚  β”‚ElastiCacheβ”‚ β”‚ DynamoDB  β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚ (Redis)  β”‚ β”‚ (Cache)   β”‚  β”‚ β”‚
              β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚ β”‚
              β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚ β”‚
              β”‚  β”‚  β”‚   SQS   β”‚ β”‚  S3 Model  β”‚  β”‚ β”‚
              β”‚  β”‚  β”‚ (Queues)β”‚ β”‚  Registry  β”‚  β”‚ β”‚
              β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚ β”‚
              β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Isolating the wrapper service, GPU compute, and data layers into separate subnets isn't just security hygiene β€” it's what lets each layer scale, fail, and be audited independently.

Recall first β€” before you read on
  • Between A100-40GB and H100-80GB, which was chosen and why?
  • What was the target p95 latency and monthly compute budget constraint?
The GPU Math: Why This MattersCAPACITY PLANNING
β–Ά
gpu-instance-decision.txtascii
INSTANCE SELECTION DECISION TREE:

Our constraints:
β”œβ”€β”€ 10,000 developers, peak concurrency ~500
β”œβ”€β”€ Target latency: p95 < 800ms
β”œβ”€β”€ Budget: $250K/month for compute
└── Must fit models in single GPU (no tensor parallelism)

OPTION 1: A100-40GB (p4d.24xlarge)
β”œβ”€β”€ 8Γ— A100 GPUs, 40GB each
β”œβ”€β”€ Can host 1 instance of DeepSeek-33B per GPU
β”œβ”€β”€ With quantization: 2 instances per GPU
β”œβ”€β”€ Per-instance cost: $32.77/hr (1-year reserved)
β”œβ”€β”€ Capacity per node: 8-16 concurrent requests
β”œβ”€β”€ 8 nodes Γ— 8 GPUs Γ— 1.5 (avg utilization) = 96 concurrent
└── COST: 8 Γ— $32.77 Γ— 730 = $191,380/month

OPTION 2: H100-80GB (p5.48xlarge)
β”œβ”€β”€ 8Γ— H100 GPUs, 80GB each
β”œβ”€β”€ 2-3Γ— faster inference than A100
β”œβ”€β”€ Per-instance cost: $98.32/hr (1-year reserved)
β”œβ”€β”€ Capacity per node: 24-32 concurrent requests
β”œβ”€β”€ 3 nodes Γ— 8 GPUs Γ— 2 = 48 concurrent (but 2Γ— faster)
└── COST: 3 Γ— $98.32 Γ— 730 = $215,200/month

WINNER: A100s
β”œβ”€β”€ Better cost/capacity ratio
β”œβ”€β”€ Sufficient performance for our latency targets
└── More flexible for fine-tuning workloads
THE REAL QUESTIONWhat happens during Monday morning peak when 40% of developers hit the system simultaneously?
What we learned

The faster, more expensive H100 wasn't the right call here β€” once you model real concurrency needs against latency targets, the 'better' chip can still be the worse business decision.

Recall first β€” before you read on
  • What's the min/max instance range for the DeepSeek endpoint, and how does it change by time of day?
  • How does the batching queue improve throughput without extra GPU memory?
Auto-Scaling ConfigurationSAGEMAKER
β–Ά
autoscaling-policy.txtascii
SAGEMAKER AUTO-SCALING POLICY:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ DeepSeek Endpoint (Primary):                       β”‚
β”‚                                                    β”‚
β”‚  Target Tracking:                                  β”‚
β”‚  β”œβ”€β”€ Metric: SageMakerVariantInvocationsPerInstanceβ”‚
β”‚  β”œβ”€β”€ Target: 8 requests per instance               β”‚
β”‚  β”œβ”€β”€ Scale-out: When > 10 for 2 minutes            β”‚
β”‚  └── Scale-in: When < 3 for 10 minutes             β”‚
β”‚                                                    β”‚
β”‚  Min Instances: 4 (off-peak)                       β”‚
β”‚  Max Instances: 16 (peak)                          β”‚
β”‚  Cool-down: 300 seconds                            β”‚
β”‚                                                    β”‚
β”‚  Scheduled Scaling:                                β”‚
β”‚  β”œβ”€β”€ Mon-Fri 8:00-10:00: Min 12 instances          β”‚
β”‚  β”œβ”€β”€ Mon-Fri 10:00-17:00: Min 8 instances          β”‚
β”‚  └── Weekends/2AM-6AM: Min 2 instances             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Batch Optimization Queue:
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                                                    β”‚
β”‚  Client sends: "Write a function to..."            β”‚
β”‚         β”‚                                          β”‚
β”‚         β–Ό                                          β”‚
β”‚  [SQS Queue accumulates for 50ms window]           β”‚
β”‚         β”‚                                          β”‚
β”‚         β–Ό                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”           β”‚
β”‚  β”‚ Batch of 4 requests formed:         β”‚           β”‚
β”‚  β”‚ [req1, req2, req3, req4]           β”‚           β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β”‚
β”‚         β”‚                                          β”‚
β”‚         β–Ό                                          β”‚
β”‚  [Single forward pass through model]               β”‚
β”‚         β”‚                                          β”‚
β”‚         β–Ό                                          β”‚
β”‚  [Responses split and returned individually]       β”‚
β”‚                                                    β”‚
β”‚  Result: 4Γ— throughput, same GPU memory            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Static capacity planning fails at this scale β€” scheduled scaling (matching known daily/weekly patterns) combined with reactive target-tracking is what kept cost and latency both in check.

04

The IDFC-Coder Wrapper: The Intelligence Layer

Recall first β€” before you read on
  • What are the three stages a request passes through before hitting the model β€” auth/rate-limit, prompt engineering, caching?
  • What two-tier caching strategy is used, and what's the combined hit-rate benefit?
Architecture Deep DiveWRAPPER SERVICE
β–Ά
wrapper-architecture.txtascii
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    IDFC-Coder Wrapper                        β”‚
β”‚                                                              β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              Request Orchestrator                     β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ Auth/      β”‚  β”‚ Rate       β”‚  β”‚ Team       β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ RBAC       │──│ Limiter    │──│ Quota Checkβ”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              Prompt Engineering Engine               β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  Template:                                            β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ System: You are a fintech code assistant.   β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ You write secure, compliant Java/Python.    β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Context from codebase:                       β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ {injected_relevant_files}                    β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Jira ticket context:                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ {ticket_description_and_acceptance_criteria} β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ User request: {user_prompt}                  β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Respond with code only, no explanations.     β”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              Caching Layer                           β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  Cache Key: MD5(prompt_hash + context_hash)          β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  L1: ElastiCache Redis (in-memory, TTL: 1 hour)      β”‚   β”‚
β”‚  β”‚  └── Hit rate: ~40% for repeated patterns            β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  L2: DynamoDB (persistent, TTL: 24 hours)            β”‚   β”‚
β”‚  β”‚  └── Hit rate: +15% for cross-instance sharing       β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  Cache Invalidation: On new model version deployment  β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              Queue & Token Management                β”‚   β”‚
β”‚  β”‚                                                       β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                  β”‚   β”‚
β”‚  β”‚  β”‚ Token Counterβ”‚  β”‚ Cost Tracker β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚              β”‚  β”‚              β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚ Input: 450   β”‚  β”‚ Team: Alpha  β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚ Output: 320  β”‚  β”‚ Daily: 2.3M  β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚ Total: 770   β”‚  β”‚ Limit: 5M    β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚ Est. Cost:   β”‚  β”‚ Remaining:   β”‚                  β”‚   β”‚
β”‚  β”‚  β”‚ $0.0008      β”‚  β”‚ 2.7M tokens  β”‚                  β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                  β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Every request the wrapper handles goes through auth, quota, prompt engineering, and caching before it ever reaches the model β€” the 'intelligence layer' is as much about governance as it is about prompting.

Recall first β€” before you read on
  • What's the effective cost per developer request after caching, and how does that compare to a commercial API?
  • How are team token budgets enforced, and what happens at 80% and 100% usage?
The Token EconomyCOST MODEL
β–Ά

This is where most self-hosted deployments fail. Let me explain why:

cost-per-inference.txtascii
COST PER INFERENCE BREAKDOWN:

DeepSeek Coder 33B on A100 (with batching):
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Fixed costs (per hour):                            β”‚
β”‚ β”œβ”€β”€ GPU instance: $32.77                          β”‚
β”‚ β”œβ”€β”€ Overhead (network, storage): $2.50            β”‚
β”‚ └── Total: $35.27/hour                            β”‚
β”‚                                                    β”‚
β”‚ Throughput at batch size 4:                        β”‚
β”‚ β”œβ”€β”€ 400 requests per GPU per hour                 β”‚
β”‚ └── Cost per request: $0.088                       β”‚
β”‚                                                    β”‚
β”‚ With caching (40% hit rate):                       β”‚
β”‚ β”œβ”€β”€ Effective cost: $0.053 per request            β”‚
β”‚ └── Compare to Claude API: $0.03-0.10              β”‚
β”‚                                                    β”‚
β”‚ BUT! For our use case:                             β”‚
β”‚ β”œβ”€β”€ Average tokens: 800 per request                β”‚
β”‚ β”œβ”€β”€ Cached requests: near-zero cost               β”‚
β”‚ └── Blended cost: $0.025 per developer request    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

TOKEN BUDGETING SYSTEM:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Team Token Allocation (Monthly):                   β”‚
β”‚                                                    β”‚
β”‚ Team Alpha (Core Platform, 200 devs):              β”‚
β”‚ β”œβ”€β”€ Budget: 50M tokens/month                      β”‚
β”‚ β”œβ”€β”€ Rate: 10K tokens/dev/day                      β”‚
β”‚ └── Alerts at 80%, blocks at 100%                 β”‚
β”‚                                                    β”‚
β”‚ Team Beta (Mobile, 150 devs):                      β”‚
β”‚ β”œβ”€β”€ Budget: 37.5M tokens/month                    β”‚
β”‚ β”œβ”€β”€ Rate: 10K tokens/dev/day                      β”‚
β”‚ └── Can request increase via Jira ticket          β”‚
β”‚                                                    β”‚
β”‚ Enforcement via API Gateway + wrapper layer        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
HARD QUESTIONWhat happens when a VP of Engineering demands unlimited access? How do you say no with data?
What we learned

Caching isn't optional at this scale β€” a 40-55% effective hit rate is what turns a marginally-competitive per-request cost into a clearly winning one, and hard token budgets are what make 'no' to unlimited access defensible with data.

05

Jira Integration: The Autonomous Agent

Recall first β€” before you read on
  • How often does the Jira agent poll, and what JQL filter does it use?
  • What are the three possible output actions once the agent generates code for a ticket?
System Design β€” Agent Orchestration FlowEVENTBRIDGE Β· 30 MIN
β–Ά
jira-agent-flow.txtascii
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    Jira Agent Controller                 β”‚
β”‚                                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              Scheduler (EventBridge)              β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚   β”‚
β”‚  β”‚  β”‚ Every 30 minutes:                          β”‚  β”‚   β”‚
β”‚  β”‚  β”‚ cron(0/30 * * * ? *)                       β”‚  β”‚   β”‚
β”‚  β”‚  β”‚                                             β”‚  β”‚   β”‚
β”‚  β”‚  β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                            β”‚  β”‚   β”‚
β”‚  β”‚  β”‚ β”‚ State Check β”‚                            β”‚  β”‚   β”‚
β”‚  β”‚  β”‚ β”‚ (DynamoDB)  β”‚                            β”‚  β”‚   β”‚
β”‚  β”‚  β”‚ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜                            β”‚  β”‚   β”‚
β”‚  β”‚  β”‚        β”‚                                    β”‚  β”‚   β”‚
β”‚  β”‚  β”‚   Last poll: 2024-01-15 14:00:00 UTC       β”‚  β”‚   β”‚
β”‚  β”‚  β”‚   Cursor: ticket-45678                      β”‚  β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚           Jira Connector (Rate Limited)          β”‚   β”‚
β”‚  β”‚                                                   β”‚   β”‚
β”‚  β”‚  API CALL STRATEGY:                               β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ JQL Query:                               β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ project = "IDFC" AND                     β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ status in ("In Development", "Ready")    β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ AND updated >= -30m                      β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ ORDER BY updated DESC                    β”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β”‚                                                   β”‚   β”‚
β”‚  β”‚  Rate Limiting:                                   β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ β€’ Jira Cloud: 10 requests/sec            β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ β€’ Our self-hosted: 50 requests/sec       β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ β€’ Token bucket algorithm:                β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   - Burst: 30                            β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   - Sustained: 10/sec                    β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ β€’ Exponential backoff: 1s, 2s, 4s, 8s   β”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚         Prompt Transformer                       β”‚   β”‚
β”‚  β”‚                                                   β”‚   β”‚
β”‚  β”‚  Ticket β†’ Prompt Mapping:                         β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ {                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "ticket_id": "IDFC-1234",              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "title": "Add fraud detection rule",   β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "description": "...",                  β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "acceptance_criteria": [               β”‚     β”‚   β”‚
β”‚  β”‚  β”‚     "Must flag transactions > $10K",     β”‚     β”‚   β”‚
β”‚  β”‚  β”‚     "Must not exceed 100ms latency"      β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   ]                                       β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ }                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                           β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Transforms to:                            β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ "Write a Java method that implements..."  β”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚                         β”‚                                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚         Result Publisher                         β”‚   β”‚
β”‚  β”‚                                                   β”‚   β”‚
β”‚  β”‚  Output Actions:                                  β”‚   β”‚
β”‚  β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚   β”‚
β”‚  β”‚  β”‚ Option A: Comment on ticket              β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ {                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "body": "πŸ€– AI Suggestion:\n```java..." β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ }                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                           β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Option B: Create subtask                 β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ {                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "summary": "Implement: ...",           β”‚     β”‚   β”‚
β”‚  β”‚  β”‚   "description": "AI-generated code..."  β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ }                                         β”‚     β”‚   β”‚
β”‚  β”‚  β”‚                                           β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ Option C: Create pull request             β”‚     β”‚   β”‚
β”‚  β”‚  β”‚ (via Bitbucket/GitHub API)                β”‚     β”‚   β”‚
β”‚  β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Turning Jira tickets into code isn't a single API call β€” it's a full pipeline of polling, rate-limited fetching, prompt transformation, and a deliberate choice of how to publish the result (comment, subtask, or PR).

Recall first β€” before you read on
  • Why can't the agent trust Jira's search index alone, and what's the dual-verification fix?
  • What's the fallback when the Jira rate limit is exhausted?
Handling Jira's QuirksCONSISTENCY
β–Ά
eventual-consistency.txtascii
EVENTUAL CONSISTENCY PATTERN:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Problem: Jira search index may be stale            β”‚
β”‚                                                    β”‚
β”‚ Solution: Dual verification                        β”‚
β”‚                                                    β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚ β”‚ 1. Get tickets from search (fast, maybe    β”‚    β”‚
β”‚ β”‚    stale)                                   β”‚    β”‚
β”‚ β”‚ 2. For each candidate, direct GET ticket   β”‚    β”‚
β”‚ β”‚    (slow, always current)                   β”‚    β”‚
β”‚ β”‚ 3. Compare versions:                        β”‚    β”‚
β”‚ β”‚    if search.version < direct.version:      β”‚    β”‚
β”‚ β”‚      Use direct version                     β”‚    β”‚
β”‚ β”‚    If already processed (idempotency key):  β”‚    β”‚
β”‚ β”‚      Skip                                    β”‚    β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β”‚                                                    β”‚
β”‚ Rate Limit Exhaustion Handling:                    β”‚
β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚ β”‚ β€’ Pre-emptive: Track X-RateLimit-Remaining β”‚    β”‚
β”‚ β”‚ β€’ Reactive: Queue overflow β†’ DLQ           β”‚    β”‚
β”‚ β”‚ β€’ Backpressure: Skip poll if previous not  β”‚    β”‚
β”‚ β”‚   complete                                  β”‚    β”‚
β”‚ β”‚ β€’ Priority queue: Urgent tickets first     β”‚    β”‚
β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
CONTROVERSIAL QUESTIONIs automatically generating code from Jira tickets actually helping developers, or are we just creating more review burden?
What we learned

Third-party SaaS APIs are eventually consistent by default β€” building a dual-verification (search + direct GET) pattern was necessary just to avoid acting on stale data.

06

Production Concerns: The Boring Stuff That Saves You

Recall first β€” before you read on
  • What failure threshold trips the circuit breaker from CLOSED to OPEN?
  • What's the retry backoff sequence before a request lands in the dead-letter queue?
Resilience PatternsCIRCUIT BREAKER
β–Ά
circuit-breaker-state-machine.txtascii
CIRCUIT BREAKER CONFIGURATION:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ State Machine:                                     β”‚
β”‚                                                    β”‚
β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                              β”‚
β”‚         β”‚  CLOSED  β”‚ (Normal operation)            β”‚
β”‚         β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜                              β”‚
β”‚              β”‚                                      β”‚
β”‚     failures >= 5 in 60s window                    β”‚
β”‚              β”‚                                      β”‚
β”‚              β–Ό                                      β”‚
β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                              β”‚
β”‚         β”‚   OPEN   β”‚ (Fail fast, return cache)     β”‚
β”‚         β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜                              β”‚
β”‚              β”‚                                      β”‚
β”‚     30 second timeout                              β”‚
β”‚              β”‚                                      β”‚
β”‚              β–Ό                                      β”‚
β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                              β”‚
β”‚         β”‚ HALF-OPENβ”‚ (Test with 1 request)         β”‚
β”‚         β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜                              β”‚
β”‚              β”‚                                      β”‚
β”‚     Success β†’ CLOSED                               β”‚
β”‚     Failure β†’ OPEN (back to timeout)               β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

RETRY STRATEGY WITH JITTER:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Attempt 1: Immediate                               β”‚
β”‚ Attempt 2: 1s + random(0-200ms)                    β”‚
β”‚ Attempt 3: 2s + random(0-400ms)                    β”‚
β”‚ Attempt 4: 4s + random(0-800ms)                    β”‚
β”‚ Attempt 5: 8s + random(0-1600ms) [MAX]             β”‚
β”‚                                                    β”‚
β”‚ Then: Dead Letter Queue                            β”‚
β”‚ └── SQS β†’ Lambda β†’ Alert β†’ Slack                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Circuit breakers and jittered retries aren't just best practice boilerplate β€” they're what stands between a slow GPU endpoint and a full outage, by failing fast and falling back to cache instead of queuing forever.

Recall first β€” before you read on
  • What canary traffic split is used for a new model version, and over what monitoring window?
  • Name two conditions that trigger an automatic rollback.
Model Deployment StrategyCANARY
β–Ά
canary-deployment.txtascii
CANARY DEPLOYMENT FOR MODEL UPDATES:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Current Model: deepseek-v1.2 (tag: prod)           β”‚
β”‚ New Model: deepseek-v1.3 (tag: canary)             β”‚
β”‚                                                    β”‚
β”‚ Traffic Splitting (SageMaker):                     β”‚
β”‚                                                    β”‚
β”‚ Endpoint: idfc-coder-prod                          β”‚
β”‚ β”œβ”€β”€ Variant-A (v1.2): 90% weight                  β”‚
β”‚ └── Variant-B (v1.3): 10% weight                  β”‚
β”‚                                                    β”‚
β”‚ Monitoring Period: 4 hours                         β”‚
β”‚                                                    β”‚
β”‚ Automatic Rollback IF:                             β”‚
β”‚ β”œβ”€β”€ p95 latency > 1200ms (baseline: 800ms)        β”‚
β”‚ β”œβ”€β”€ Error rate > 5% (baseline: 1%)                β”‚
β”‚ β”œβ”€β”€ Token output length < 50 (hallucination risk) β”‚
β”‚ └── User rejection rate > 20%                     β”‚
β”‚                                                    β”‚
β”‚ Gradual Promotion:                                 β”‚
β”‚ β”œβ”€β”€ Hour 4: 25% traffic                            β”‚
β”‚ β”œβ”€β”€ Hour 8: 50% traffic                            β”‚
β”‚ β”œβ”€β”€ Hour 12: 100% traffic                          β”‚
β”‚ └── After 24h stable: decommission v1.2            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Model updates need the same production rigor as code deploys β€” canary traffic splitting with automatic rollback thresholds (latency, error rate, even output length as a hallucination signal) catches regressions before they reach everyone.

Recall first β€” before you read on
  • What's the split between business metrics and technical metrics on the dashboard?
  • In the X-Ray trace example, which stage consumed the most time?
Observability StackCLOUDWATCH / X-RAY
β–Ά
observability-stack.txtascii
MONITORING ARCHITECTURE:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              CloudWatch Dashboard                   β”‚
β”‚                                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚ Business Metrics  β”‚  β”‚ Technical Metrics β”‚       β”‚
β”‚  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€       β”‚
β”‚  β”‚ β€’ Requests/team   β”‚  β”‚ β€’ p50/p95/p99 lat β”‚       β”‚
β”‚  β”‚ β€’ Tokens consumed  β”‚  β”‚ β€’ Queue depth     β”‚       β”‚
β”‚  β”‚ β€’ Cache hit rate  β”‚  β”‚ β€’ GPU utilization β”‚       β”‚
β”‚  β”‚ β€’ Cost per team   β”‚  β”‚ β€’ Error rates     β”‚       β”‚
β”‚  β”‚ β€’ Code accepted%  β”‚  β”‚ β€’ Circuit breaks  β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β”‚                                                    β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚  β”‚              X-Ray Tracing                  β”‚    β”‚
β”‚  β”‚                                              β”‚    β”‚
β”‚  β”‚  Trace: request-abc123                       β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ API Gateway: 12ms                       β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ Auth check: 8ms                         β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ Cache lookup: 3ms [MISS]                β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ SQS Queue time: 45ms                    β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ LLM Inference: 340ms                    β”‚    β”‚
β”‚  β”‚  β”œβ”€β”€ Response transform: 5ms                 β”‚    β”‚
β”‚  β”‚  └── TOTAL: 413ms                            β”‚    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Business metrics (cost, adoption, code-accepted%) and technical metrics (latency, GPU utilization) have to sit on the same dashboard β€” otherwise engineering optimizes for speed while the business case quietly erodes.

Recall first β€” before you read on
  • Soft or hard multi-tenancy β€” which was chosen, and what's the main cost trade-off with the other option?
  • What's the biggest risk of the soft multi-tenancy model?
Multi-Tenancy IsolationSOFT vs HARD
β–Ά
tenant-isolation.txtascii
TENANT ISOLATION MODEL:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Option A: Soft Multi-tenancy (We chose this)       β”‚
β”‚                                                    β”‚
β”‚ β€’ Shared inference endpoints                       β”‚
β”‚ β€’ Queue-level isolation (per-team SQS queues)      β”‚
β”‚ β€’ Token quotas enforced at wrapper layer           β”‚
β”‚ β€’ Separate DynamoDB tables per team for caching    β”‚
β”‚ β€’ CloudWatch metrics tagged with TeamID            β”‚
β”‚                                                    β”‚
β”‚ Pros: Cost effective, simple operations            β”‚
β”‚ Cons: Noisy neighbor possible                      β”‚
β”‚                                                    β”‚
β”‚ Option B: Hard Multi-tenancy (Rejected)            β”‚
β”‚                                                    β”‚
β”‚ β€’ Separate SageMaker endpoints per team            β”‚
β”‚ β€’ Full VPC isolation                                β”‚
β”‚ β€’ Pros: True isolation                              β”‚
β”‚ β€’ Cons: 50 teams Γ— 2 GPUs = $1.6M/month minimum   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
STRATEGIC QUESTIONIn fintech, regulators might want hard multi-tenancy. How do we prepare for that without over-engineering now?
What we learned

Soft multi-tenancy was the pragmatic choice, not the ideal one β€” the team consciously accepted 'noisy neighbor' risk to avoid a 10x cost blowout, while leaving a path open to harden isolation later if regulators demand it.

07

Iterative Challenges: The War Stories

Recall first β€” before you read on
  • What was the p95 latency before optimization, and what were the main root causes?
  • Which single change β€” batching, streaming, quantization, or ECS β€” do you think mattered most, and what was the final p95?
Challenge 1 β€” The Latency NightmareP1 INCIDENT
β–Ά
latency-phase1.txtascii
PHASE 1 - Unoptimized (Week 1):

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Single request flow:                                β”‚
β”‚ Client β†’ API Gateway β†’ Lambda β†’ SageMaker β†’ Client β”‚
β”‚                                                    β”‚
β”‚ Results:                                           β”‚
β”‚ β”œβ”€β”€ p50: 1200ms                                    β”‚
β”‚ β”œβ”€β”€ p95: 3500ms                                    β”‚
β”‚ └── User feedback: "Slower than Copilot!"          β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Root cause analysis β€” why so slow?

root-cause.txtascii
Why so slow?
β”œβ”€β”€ No batching: One forward pass per request
β”œβ”€β”€ Cold starts: Lambda initialization 200-400ms
β”œβ”€β”€ No streaming: Waiting for full response
β”œβ”€β”€ Model too big: 33B parameters, lots of compute
└── Network latency: Cross-AZ calls
latency-phase2.txtascii
PHASE 2 - Optimized (Week 3):

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Optimized flow:                                     β”‚
β”‚ Client β†’ ALB β†’ ECS (warm) β†’ Batched β†’ Streaming   β”‚
β”‚                                                    β”‚
β”‚ Changes made:                                      β”‚
β”‚ β”œβ”€β”€ Switched to ECS Fargate (no cold starts)       β”‚
β”‚ β”œβ”€β”€ Dynamic batching (50ms accumulation window)    β”‚
β”‚ β”œβ”€β”€ Token streaming via Server-Sent Events         β”‚
β”‚ β”œβ”€β”€ Model quantization: FP32 β†’ INT8 (2Γ— speed)     β”‚
β”‚ β”œβ”€β”€ KV-cache optimization for repeated prefixes    β”‚
β”‚ └── Cross-AZ placement groups for SageMaker        β”‚
β”‚                                                    β”‚
β”‚ Results:                                           β”‚
β”‚ β”œβ”€β”€ p50: 340ms (3.5Γ— improvement)                  β”‚
β”‚ β”œβ”€β”€ p95: 720ms (4.8Γ— improvement)                  β”‚
β”‚ └── Throughput: 800 req/sec (from 200)             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

The first version's 3.5s p95 wasn't a model problem β€” it was an architecture problem (cold starts, no batching, no streaming). Fixing the plumbing delivered a bigger win than any model swap would have.

Recall first β€” before you read on
  • Why did the naive FP16 memory math not even fit on a 40GB A100?
  • What three techniques finally brought memory usage under budget?
Challenge 2 β€” The OOM CrisisOUTAGE
β–Ά

What happened: Tuesday 14:00, production outage β€” SageMaker returning 500s. Root cause: DeepSeek-33B OOM on the 40GB A100 because the context window grew to 16K tokens.

memory-analysis.txtascii
MEMORY ANALYSIS:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ GPU Memory Budget (40GB):                          β”‚
β”‚                                                    β”‚
β”‚ Model weights (FP16): 33B Γ— 2 bytes = 66GB         β”‚
β”‚                                                    β”‚
β”‚ Wait, that doesn't fit!                             β”‚
β”‚                                                    β”‚
β”‚ Let's recalculate with quantization:               β”‚
β”‚                                                    β”‚
β”‚ Model weights (INT8): 33B Γ— 1 byte = 33GB         β”‚
β”‚ KV Cache (batch_size=4, seq_len=4096): 8GB         β”‚
β”‚ Activations: 4GB                                    β”‚
β”‚ CUDA overhead: 2GB                                  β”‚
β”‚ ─────────────────────────────────────               β”‚
β”‚ Total: 47GB ❌ STILL DOESN'T FIT!                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
the-fix.txtascii
THE FIX:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Solution 1: Flash Attention 2                       β”‚
β”‚ β”œβ”€β”€ Reduces memory by 5-10Γ— for attention          β”‚
β”‚ └── KV Cache: 8GB β†’ 1.5GB                          β”‚
β”‚                                                    β”‚
β”‚ Solution 2: Gradient checkpointing (even for inf)   β”‚
β”‚ └── Activations: 4GB β†’ 1GB                         β”‚
β”‚                                                    β”‚
β”‚ Solution 3: Batch size reduction                    β”‚
β”‚ └── From 4 to 2 during peak                        β”‚
β”‚                                                    β”‚
β”‚ New Memory Profile:                                β”‚
β”‚ β”œβ”€β”€ Weights: 33GB                                  β”‚
β”‚ β”œβ”€β”€ KV Cache: 1.5GB                                β”‚
β”‚ β”œβ”€β”€ Activations: 1GB                               β”‚
β”‚ β”œβ”€β”€ Overhead: 2GB                                   β”‚
β”‚ └── Total: 37.5GB βœ… FITS!                         β”‚
β”‚                                                    β”‚
β”‚ Also: Added automatic context window adjustment    β”‚
β”‚ β”œβ”€β”€ Monitor: GPU memory utilization                β”‚
β”‚ β”œβ”€β”€ If > 85%: Reduce max_tokens by 25%             β”‚
β”‚ └── Alert: If sustained > 90% for 5 min           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Back-of-envelope memory math (FP16 weights alone) is dangerously misleading β€” real GPU budgets have to account for KV cache, activations, and CUDA overhead together, and quantization plus Flash Attention were what finally made the numbers fit.

Recall first β€” before you read on
  • What are the three types of hallucination encountered β€” think APIs, business logic, and security?
  • What does the model do instead of guessing when it's unsure about a policy or API?
Challenge 3 β€” Hallucination MitigationQUALITY
β–Ά

Hallucination types encountered:

  1. Non-existent APIs β€” model invents calls like BankAPI.getInstance().getV2Client(), an internal API not present in training data.
  2. Incorrect business logic β€” model suggests a $10,000 fraud threshold when policy requires $5,000.
  3. Security anti-patterns β€” string-concatenated SQL, a firing offense in fintech code.
Solution β€” Grounded Prompt Engineering β–Ά
grounded-system-prompt.txtascii
Enhanced System Prompt:

You are IDFC-First Bank's internal code assistant.

CONSTRAINTS (enforced):
1. Only use APIs from our SDK:
   com.idfc.banking.v3.*
   com.idfc.fraud.detection.*
   com.idfc.compliance.*

2. All database queries MUST use parameterized
   statements (PreparedStatement only)

3. All monetary values MUST use BigDecimal,
   never float/double

4. All external calls MUST have:
   - Timeout (max 30s)
   - Circuit breaker
   - Retry with backoff

5. If you're unsure about an API or policy,
   respond with: "NEEDS_CLARIFICATION: [question]"
   Do not guess.

CONTEXT INJECTION:
{relevant_code_from_codebase}
{relevant_confluence_docs}
{jira_acceptance_criteria}
Post-processing Validation β–Ά
post-processing.txtascii
Before returning to user:

1. Static Analysis (SonarQube rules):
   β”œβ”€β”€ No SQL injection patterns
   β”œβ”€β”€ No hardcoded credentials
   └── Checked exceptions handled

2. API Validation:
   β”œβ”€β”€ Verify all method calls exist in SDK
   └── If not found: replace with comment

3. Policy Check:
   β”œβ”€β”€ Fraud thresholds from config service
   └── Inject correct values
What we learned

Prompting alone can't be trusted to stop hallucinated APIs or wrong compliance thresholds β€” it took explicit constraints, a 'don't guess, ask' instruction, and a post-processing validation layer to make outputs safe to ship.

Recall first β€” before you read on
  • What are the four adoption phases, from pioneers to laggards?
  • How was the objection 'it doesn't understand our codebase' addressed?
Challenge 4 β€” The 10,000 User RolloutADOPTION
β–Ά
phased-adoption.txtascii
PHASED ADOPTION STRATEGY:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Phase 1: Pioneers (Month 1, 100 users)             β”‚
β”‚ β”œβ”€β”€ Selected power users from Core Platform team   β”‚
β”‚ β”œβ”€β”€ Direct Slack channel with platform team        β”‚
β”‚ β”œβ”€β”€ Weekly feedback sessions                       β”‚
β”‚ └── Goal: Iron out UX issues, gather testimonials  β”‚
β”‚                                                    β”‚
β”‚ Phase 2: Early Adopters (Month 2, 1000 users)      β”‚
β”‚ β”œβ”€β”€ Add 2-3 more teams                             β”‚
β”‚ β”œβ”€β”€ Introduce IDE plugin (VS Code, IntelliJ)       β”‚
β”‚ β”œβ”€β”€ Gamification: Leaderboard of tokens saved      β”‚
β”‚ └── Goal: Prove productivity improvement           β”‚
β”‚                                                    β”‚
β”‚ Phase 3: Majority (Month 3-4, 5000 users)          β”‚
β”‚ β”œβ”€β”€ Self-service onboarding portal                 β”‚
β”‚ β”œβ”€β”€ Required: 30-min training video                β”‚
β”‚ β”œβ”€β”€ Auto-enrollment for new hires                  β”‚
β”‚ └── Goal: Make it default, not optional            β”‚
β”‚                                                    β”‚
β”‚ Phase 4: Laggards (Month 5-6, 10000 users)         β”‚
β”‚ β”œβ”€β”€ Mandatory for certain ticket types             β”‚
β”‚ β”œβ”€β”€ "AI-first" workflow: Must review AI suggestion  β”‚
β”‚ β”‚   before writing from scratch                    β”‚
β”‚ └── Goal: Complete transformation                  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
adoption-friction.txtascii
ADOPTION FRICTION POINTS & SOLUTIONS:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ "It's slower than Copilot"                         β”‚
β”‚ β†’ Show latency metrics, it's actually faster for   β”‚
β”‚   our codebase because of context injection        β”‚
β”‚                                                    β”‚
β”‚ "The code is wrong"                                β”‚
β”‚ β†’ Add "Report bad suggestion" button               β”‚
β”‚ β†’ Track rejection rate, feed back to fine-tuning   β”‚
β”‚ β†’ Public dashboard showing improvement over time   β”‚
β”‚                                                    β”‚
β”‚ "I don't trust AI"                                 β”‚
β”‚ β†’ Start with test generation (low risk)            β”‚
β”‚ β†’ Show examples of bugs caught by AI suggestions   β”‚
β”‚ β†’ Peer pressure: "Your team's top performer uses it"β”‚
β”‚                                                    β”‚
β”‚ "It doesn't understand our codebase"               β”‚
β”‚ β†’ RAG pipeline: Index all repos, retrieve context  β”‚
β”‚ β†’ Fine-tune on our private code (with permissions) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Technology readiness and organizational readiness are two different problems β€” the phased pioneer β†’ early adopter β†’ majority β†’ laggard rollout mattered as much as any latency or accuracy fix.

08

Concrete Metrics: The Scorecard

Six-month post-launch results.

54.4%
Year-1 cost reduction
$4.28M
Year-1 savings
720ms
p95 latency, simple gen
1,200 rps
Peak throughput
68%
Acceptance rate, Month 6
+42 NPS
Developer sentiment
Recall first β€” before you read on
  • What was the actual realized savings percentage compared to the original projection?
Cost AnalysisFINANCE
β–Ά
cost-analysis.txtascii
COST ANALYSIS:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Previous Year (Copilot Enterprise + Claude API):       β”‚
β”‚ β”œβ”€β”€ GitHub Copilot: 10,000 seats Γ— $39/month          β”‚
β”‚ β”‚   = $4,680,000/year                                  β”‚
β”‚ β”œβ”€β”€ Anthropic API (heavy users): $3,200,000/year       β”‚
β”‚ └── TOTAL: $7,880,000/year                             β”‚
β”‚                                                        β”‚
β”‚ Self-Hosted Solution (Year 1):                         β”‚
β”‚ β”œβ”€β”€ Compute (A100 reserved): $2,296,560                β”‚
β”‚ β”œβ”€β”€ Storage & Networking: $300,000                     β”‚
β”‚ β”œβ”€β”€ Engineering Team (4 FTE): $800,000                 β”‚
β”‚ β”œβ”€β”€ Training & Onboarding: $200,000                    β”‚
β”‚ └── TOTAL: $3,596,560                                  β”‚
β”‚                                                        β”‚
β”‚ SAVINGS: $4,283,440 (54.4% reduction)                   β”‚
β”‚                                                        β”‚
β”‚ Year 2 Projection (with optimization):                 β”‚
β”‚ β”œβ”€β”€ Better batching, higher utilization: -15% compute  β”‚
β”‚ β”œβ”€β”€ More caching, smaller models: -10%                 β”‚
β”‚ └── Projected: $2,900,000 (63.2% savings)             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
What we learned

Real-world savings rarely match the initial projection exactly β€” tracking actuals against the original business case is what keeps the initiative honest and fundable.

Recall first β€” before you read on
  • How much did p95 latency improve from the unoptimized to the optimized state?
Performance MetricsLATENCY / THROUGHPUT
β–Ά
performance-metrics.txtascii
Latency (DeepSeek Coder 33B, INT8):

Simple completions (< 200 tokens):
β”œβ”€β”€ p50: 280ms
β”œβ”€β”€ p95: 620ms
β”œβ”€β”€ p99: 980ms
└── SLA: p95 < 800ms βœ…

Complex generations (200-1000 tokens):
β”œβ”€β”€ p50: 450ms
β”œβ”€β”€ p95: 890ms
β”œβ”€β”€ p99: 1400ms
└── Streaming: First token in < 100ms

Throughput:
β”œβ”€β”€ Peak: 1,200 requests/second
β”œβ”€β”€ Average: 450 requests/second
β”œβ”€β”€ Daily volume: ~15M requests
└── GPU utilization: 73% average
What we learned

The multi-week latency optimization effort translated directly into a metric leadership actually cares about β€” this is the proof point that turns an engineering win into a business one.

Recall first β€” before you read on
  • What direction did bugs-per-release and security vulnerabilities move after rollout?
Quality MetricsCORRECTNESS
β–Ά
quality-metrics.txtascii
Internal Benchmark (vs. Human Baseline):

Code Correctness (passes tests first try):
β”œβ”€β”€ Simple CRUD: 91% (human: 95%)
β”œβ”€β”€ Complex business logic: 72% (human: 88%)
β”œβ”€β”€ SQL queries: 88% (human: 93%)
└── Unit tests: 76% (human: 90%)

Developer Acceptance Rate:
β”œβ”€β”€ Month 1: 34% (rejected or heavily modified)
β”œβ”€β”€ Month 3: 52%
β”œβ”€β”€ Month 6: 68%
└── Target: 80% by Month 12

Time Savings:
β”œβ”€β”€ Boilerplate code: 85% faster
β”œβ”€β”€ API integrations: 60% faster
β”œβ”€β”€ Bug fixes: 45% faster
└── Average across all tasks: 55% faster
What we learned

Fewer bugs and fewer vulnerabilities post-rollout suggest the guardrails (grounded prompting, post-processing validation) were doing real work, not just adding friction.

Recall first β€” before you read on
  • Roughly how did user adoption progress from Phase 1 pioneers to full rollout?
Adoption CurveGROWTH
β–Ά
adoption-curve.txtascii
Monthly Active Users:

Month 1: 87 (87% of Phase 1)
Month 2: 845 (84.5% of Phase 2)
Month 3: 3,200 (64% of Phase 3 target)
Month 4: 4,890 (97.8% of Phase 3 target)
Month 5: 7,200 (72% of total)
Month 6: 8,900 (89% of total)

Daily Active Users: ~5,200 (52%)
Average sessions per user: 3.4 per day
Average requests per user: 47 per day

Net Promoter Score: +42 (considered "Great")
What we learned

Adoption doesn't happen by mandate alone β€” the curve reflects the deliberate phased strategy, where trust was earned with power users before it was ever required of everyone.

Recall first β€” before you read on
  • What's one concrete velocity or business metric that improved post-adoption?
Business ImpactVELOCITY
β–Ά
business-impact.txtascii
Development Velocity:

Story Points Completed per Sprint:
β”œβ”€β”€ Before: 1,200 points (avg across teams)
β”œβ”€β”€ After 3 months: 1,450 points (+20.8%)
└── After 6 months: 1,680 points (+40%)

Time-to-Market:
β”œβ”€β”€ Feature delivery: 22 days β†’ 15 days (-31.8%)
└── Hot fixes: 4.2 hours β†’ 2.1 hours (-50%)

Code Quality:
β”œβ”€β”€ Bugs per release: 45 β†’ 38 (-15.6%)
β”œβ”€β”€ Security vulnerabilities: 12/month β†’ 8/month
└── Code review comments: +23% (AI catches patterns)
What we learned

The ultimate test of any platform investment is whether it shows up in delivery velocity, not just in usage dashboards.

09

CTO's Final Thoughts

Recall first β€” before you read on
  • What was the deliberate choice about which team to onboard first, and why?
What We Got RightRETROSPECTIVE
β–Ά
what-we-got-right.txtascii
1. STARTED WITH THE HARDEST TEAM
   └── Core Platform team was most demanding, most skeptical
   └── If we could satisfy them, others would follow

2. BUILT FOR FAILURE
   └── Every component has a fallback
   └── Circuit breakers everywhere
   └── "What happens when this breaks at 3 AM?"

3. MEASURED OBSESSIVELY
   └── You can't improve what you don't measure
   └── Every decision backed by data
   └── Killed features that didn't show impact

4. TREATED IT AS PRODUCT, NOT TOOL
   └── UX design, not just API
   └── Developer advocacy team
   └── Regular town halls for feedback
What we learned

Starting with the most skeptical, most demanding team was a deliberate stress-test β€” if the platform could win them over, broader rollout became a much easier sell.

Recall first β€” before you read on
  • What's the biggest regret around initial model size and scope?
  • What GPU capacity risk should have been planned for from day one?
What We'd Do DifferentlyLESSONS
β–Ά
what-wed-do-differently.txtascii
1. START SMALLER
   └── Begin with 7B model for autocomplete only
   └── Add complexity gradually
   └── We overbuilt the initial system

2. INVEST IN FINE-TUNING EARLIER
   └── Generic models hallucinate on proprietary APIs
   └── Fine-tuning on our codebase would have saved months
   └── Now running weekly fine-tuning jobs

3. PLAN FOR GPU SHORTAGES
   └── A100 availability was constrained
   └── Should have reserved instances from day 1
   └── Consider multi-cloud (GCP TPUs as backup)

4. OVER-COMMUNICATE COSTS
   └── Developers don't see GPU bills
   └── Should show "this request cost $0.003"
   └── Creates natural conservation behavior
What we learned

Overbuilding early and underestimating GPU supply constraints were the costliest mistakes β€” starting smaller and reserving capacity earlier would have saved months.

Recall first β€” before you read on
  • Are we making developers better, or just faster β€” what's the concern behind that question?
  • What safeguards were added to counter AI-generated 'code debt'?
The Real Question Nobody AsksREFLECTION
β–Ά
"ARE WE ACTUALLY MAKING DEVELOPERS BETTER, OR JUST FASTER?" This is the uncomfortable truth: AI code generation can create "code debt" at unprecedented speed.

We're now investing in:

  • AI-generated code review tools (fighting fire with fire)
  • Mandatory test coverage thresholds for AI code
  • "Explain the code" sessions in PR reviews
  • Tracking long-term maintenance burden of AI-generated code

The goal isn't to replace developers β€” it's to remove the boring parts so they can focus on architecture, security, and innovation. But we must be vigilant that we're not creating a generation of developers who can't code without AI assistance.

What we learned

Speed isn't the same as skill β€” the team had to deliberately invest in code review discipline and maintainability tracking to make sure AI-assisted velocity didn't quietly become AI-generated technical debt.

10

Appendix: Key Configuration Files

Recall first β€” before you read on
  • What instance type and initial instance count does the production endpoint config use?
SageMaker Deployment ConfigYAML
β–Ά
deepseek-coder-endpoint.yamlyaml
# deepseek-coder-endpoint.yaml
EndpointConfigName: idfc-coder-v1-3
ProductionVariants:
  - VariantName: primary
    ModelName: deepseek-coder-33b-instruct-int8
    InstanceType: ml.p4d.24xlarge
    InitialInstanceCount: 4
    InitialVariantWeight: 1.0
    
    # Auto-scaling
    AutoScalingPolicy:
      TargetTrackingScalingPolicyConfiguration:
        TargetValue: 8.0  # Requests per instance
        ScaleInCooldown: 600
        ScaleOutCooldown: 120
        CustomizedMetricSpecification:
          MetricName: ApproximateBacklogSizePerInstance
          Namespace: AWS/SageMaker
          Statistic: Average
What we learned

Auto-scaling policy lives in config, not just documentation β€” the endpoint config is the actual source of truth for how the system behaves under load.

Recall first β€” before you read on
  • What are the default vs fast model names, and what's the Redis cache TTL?
Wrapper Service ConfigurationPYTHON
β–Ά
idfc_coder_config.pypython
# idfc_coder_config.py
class IDFCCoderConfig:
    # Model selection
    DEFAULT_MODEL = "deepseek-coder-33b"
    FAST_MODEL = "mistral-7b-finetuned"
    
    # Batching
    MAX_BATCH_SIZE = 4
    BATCHING_WINDOW_MS = 50
    
    # Token limits
    MAX_INPUT_TOKENS = 8000
    MAX_OUTPUT_TOKENS = 2000
    COST_PER_1K_TOKENS = 0.002  # USD
    
    # Team quotas (monthly tokens)
    TEAM_QUOTAS = {
        "core-platform": 50_000_000,
        "mobile": 37_500_000,
        "web": 30_000_000,
        "data-engineering": 25_000_000,
    }
    
    # Cache
    REDIS_TTL_SECONDS = 3600  # 1 hour
    DYNAMODB_TTL_HOURS = 24
    
    # Circuit breaker
    CIRCUIT_BREAKER_THRESHOLD = 5  # failures
    CIRCUIT_BREAKER_TIMEOUT = 30   # seconds
    
    # Jira
    JIRA_POLL_INTERVAL_MINUTES = 30
    JIRA_MAX_TICKETS_PER_POLL = 50
What we learned

Team quotas, cache TTLs, and circuit-breaker thresholds are all just numbers in a config file β€” but they're the numbers that encode every governance and cost decision made earlier in this document.

This architecture represents 18 months of evolution, countless production incidents, and a fundamental shift in how our organization writes software.
SELF-HOSTING AI ISN'T JUST ABOUT COST β€” IT'S ABOUT SOVEREIGNTY OVER YOUR ENGINEERING FUTURE.