From first principles.
To production AI.
Learn → build → debug → explain → optimize → deploy. One deliberate path through AI and GenAI engineering.
Progress reflects the original roadmap snapshot. It is a learning plan, not a claim of completed expertise.
00Phase 0 — Foundations
Python · Git/GitHub · SQL · Machine Learning
COMPLETED+
Completed
- Python programming and practical software work
- Git, GitHub, virtual environments and project structure
- SQL querying, joins, aggregation and problem solving
- ML fundamentals and supervised / unsupervised learning
- scikit-learn workflow and evaluation metrics
Do not restart these foundations. Revisit only when a project or interview requires them.
01Phase 1 — Neural Network Foundations
Neuron → forward pass → loss → gradients → backpropagation
COMPLETED+
Learned + implemented
- Inputs, weights, bias and weighted sum
- Forward pass and prediction
- Loss and prediction error
- Gradient descent and parameter updates
- Derivatives and chain rule
- Backpropagation and gradient accumulation
The mechanics were implemented from scratch before relying on frameworks.
02Phase 2 — Activation Functions
Step · Sigmoid · Tanh · ReLU · Leaky ReLU · Softmax
CURRENT+
Current sequence
- Step function — basic understanding
- Sigmoid — implementation and purpose
- Tanh — intuition and use
- ReLU — importance and practical behavior
- Leaky ReLU — understand the problem it addresses
- Softmax — classification output and probability distribution
For every activation: what, why, advantage, problem, where used, tiny implementation.
No research-level derivations. Finish the engineering understanding and move on.
03Phase 3 — Multi-Layer Neural Network
Build a complete network from scratch
NEXT+
Build
- Input and hidden layers
- Multiple neurons and layers
- Output layer
- Forward propagation
- Activation and loss
- Backpropagation and parameter updates
- Mini-batch concept
Small multi-layer classifier without NumPy/PyTorch first.
04Phase 4 — NumPy for AI
Vectors · matrices · batches · vectorization
UP NEXT+
Core
- Arrays, shapes and dimensions
- Indexing and reshaping
- Broadcasting
- Vectorization
- Matrix multiplication and dot products
- Batch operations
Convert the neural network from loop-based operations into matrix-based operations.
05Phase 5 — PyTorch
Tensors · autograd · models · datasets · training loops
MUST KNOW+
Core
- Tensor operations and shapes
- Dataset and DataLoader
- nn.Module and forward()
- Loss functions and optimizers
- Autograd
- Training and validation loops
- Metrics, saving and loading models
- Inference
Data → Dataset → DataLoader → Model → Forward → Loss → Backward → Optimizer → Validation → Save.
06Phase 6 — CPU + GPU Engineering
CUDA · memory · throughput · benchmarking
MUST KNOW+
DGX Spark lab
- CPU vs GPU
- CUDA and device selection
- Move models and tensors to GPU
- GPU memory and batch size
- GPU utilization
- Training and inference time
- Tokens/sec and throughput
Run the same model on CPU and GPU and document training time, inference time, memory and throughput.
07Phase 7 — Computer Vision Support
CNN fundamentals without turning CV into a second curriculum
SUPPORTING+
Learn
- Convolution, kernels, feature maps
- Stride and padding
- Pooling
- CNN architecture and training
- Transfer learning and pretrained models
- Data augmentation, dropout and batch normalization
- OpenCV basics
Advanced CV architectures, research-level vision and multiple CV projects are optional.
08Phase 8 — NLP Foundations
Text representation → embeddings → sequence concepts
FOUNDATION+
Learn
- Text cleaning and tokenization
- Vocabulary, padding and truncation
- Special tokens
- Word embeddings and vector representation
- Cosine similarity
- RNN, LSTM and GRU concepts
- Vanishing gradient intuition
RNN/LSTM are conceptual foundations. Do not spend weeks mastering them.
09Phase 9 — Transformers
Attention · self-attention · encoder/decoder · architecture
CORE+
Must know
- Why Transformers
- Query, Key and Value
- Attention scores
- Scaled dot-product attention
- Multi-head attention
- Positional encoding
- Feed-forward networks
- Residual connections and layer normalization
- Encoder, decoder and architecture differences
- BERT vs GPT
Implement a small self-attention mechanism, then move to PyTorch/Hugging Face.
Once attention and Transformer architecture can be explained and implemented, avoid research-level mathematics.
10Phase 10 — Hugging Face
Transformers · tokenizers · pretrained models · inference
CORE+
Learn
- Transformers ecosystem
- Tokenizers
- Pretrained models
- Pipelines
- AutoTokenizer and AutoModel
- Datasets
- Model configuration and model cards
Run BERT/DistilBERT → tokenize → inference → evaluate.
11Phase 11 — Local LLM Engineering
LLM internals · generation · local inference · benchmarking
GENAI CORE+
Must know
- Decoder-only LLMs and GPT-style models
- Parameters and context window
- Tokens, IDs, embeddings and logits
- Next-token prediction
- Temperature, top-k and top-p
- Sampling and inference
- KV cache — conceptual understanding
Run open-source LLMs locally and measure model size, context, memory, latency and tokens/sec.
Publish an LLM benchmarking report.
12Phase 12 — Embeddings + Vector Search
Dense vectors · similarity · top-k · vector databases
CORE+
Learn
- Embedding models
- Dense vector representations
- Semantic similarity
- Cosine similarity
- Chunk and query embeddings
- Vector database concepts
- Top-k retrieval and metadata filtering
FAISS, Chroma or Qdrant. Do not learn every vector database.
13Phase 13 — RAG
Ingestion → chunking → embeddings → retrieval → grounded generation
CORE JOB SKILL+
Must know
- Document loading and chunking
- Chunk size and overlap
- Embeddings and vector storage
- Retrieval and top-k
- Metadata filtering
- Prompt construction and context injection
- RAG evaluation and hallucination
- Retrieval failure analysis
Production-style RAG assistant with upload, embeddings, vector DB, local LLM, citations, evaluation and API.
14Phase 14 — LangChain + LlamaIndex
Frameworks for models, prompts, retrieval and orchestration
TOOLS+
Learn
- Models and prompts
- Chains and retrievers
- Tools and structured output
- Document ingestion
- Indexing and query engines
Understand the underlying architecture first. Frameworks are implementation tools, not the foundation.
15Phase 15 — Function Calling + Tools
Schemas · tool execution · structured arguments · failures
CORE+
Build
- Tool calling and function schemas
- Structured arguments
- Tool execution and results
- Error handling
- Tool selection and permissions
- Calculator, search, file and database tools
Turn an LLM from a text generator into a system that can interact with external capabilities.
16Phase 16 — AI Agents
Reason → tool → observe → reason → act
CORE+
Must know
- Agent concept
- ReAct
- Tool use
- Planning
- Observation
- State and memory
- Multi-step execution
- Failure handling
Build a research agent that searches, reads, calculates, summarizes and answers.
17Phase 17 — LangGraph
Stateful graphs · routing · loops · human-in-the-loop
ORCHESTRATION+
Learn
- Nodes and edges
- State
- Conditional routing
- Loops
- Human-in-the-loop
- Checkpointing
- Agent orchestration
Research → Analyze → Write → Review → Improve → Final Answer.
18Phase 18 — Multi-Agent Systems
Supervisor · router · parallel agents · shared state
ADVANCED+
Learn
- Agent roles and communication
- Shared state
- Supervisor pattern
- Router pattern
- Hub-and-spoke architecture
- Parallel agents
- Failure handling
Research Agent + Writer Agent + Reviewer Agent + Supervisor.
19Phase 19 — MCP
Servers · clients · tools · resources · discovery
MODERN AI+
Must know
- MCP architecture
- Client and server
- Tools
- Resources
- Prompts
- Tool discovery
- Tool execution
Custom MCP server exposing files, databases and custom APIs to a local LLM/agent.
20Phase 20 — LoRA / QLoRA / PEFT
Parameter-efficient fine-tuning on DGX Spark
DGX CORE+
Must know
- Why fine-tuning
- Pretraining vs fine-tuning
- PEFT
- LoRA and QLoRA
- Adapters and rank
- Trainable parameters
- Dataset preparation
- Training configuration and evaluation
Fine-tune a domain LLM with LoRA/QLoRA, evaluate it, quantize it and serve it.
21Phase 21 — Quantization + Inference Optimization
Precision · memory · latency · throughput
OPTIMIZATION+
Learn
- Why quantization
- FP32, FP16/BF16, INT8 and 4-bit
- Memory and accuracy tradeoffs
- Latency and throughput
- Batch size
- KV cache
- Quantized inference
Compare precision modes using memory, latency, tokens/sec and quality.
22Phase 22 — LLM Serving
FastAPI · streaming · batching · model servers
PRODUCTION+
Learn
- FastAPI and REST APIs
- Request/response design
- Streaming responses
- Async basics
- Batching and concurrency
- GPU memory
- Latency and throughput
- vLLM / TGI / Ollama concepts
Client → API → Model Server → GPU → Response.
23Phase 23 — Docker + MLOps / LLMOps
Containers · tracking · logging · monitoring · reproducibility
PRODUCTION+
Learn
- Docker images and containers
- Dockerfile, volumes and networks
- Environment variables
- Docker Compose basics
- Experiment tracking
- Model versioning
- Logging and monitoring
- Evaluation and reproducibility
Learn MLflow enough for practical experiment tracking and model workflows.
24Phase 24 — Cloud
Compute · storage · IAM · containers · GPU deployment
SUPPORTING+
AWS fundamentals
- Cloud compute
- Storage
- IAM
- Networking basics
- Containers
- GPU instances
- EC2
- S3
Cloud supports the AI engineering path. Certification prep must not consume core GenAI time.
25Phase 25 — AI System Design
Architecture · scalability · tradeoffs · production decisions
INTERVIEW CORE+
Design
- Production RAG systems
- Agent systems
- Model serving architecture
- API gateways and application layers
- Vector DB and database integration
- Monitoring
- Latency, cost and quality tradeoffs
- RAG vs fine-tuning
- Local model vs API
Be able to draw and explain a production AI architecture from client to GPU and back.
26Phase 26 — AI Evaluation
Measure retrieval, generation, agents and production behavior
MUST KNOW+
Evaluate
- Retrieval accuracy and context relevance
- Answer correctness and faithfulness
- Hallucination rate
- Tool selection and task success
- Agent failure rate
- Latency and tokens/sec
- Memory and cost
Do not build AI systems without a way to measure whether they work.
27Phase 27 — Flagship Production AI System
RAG + Agents + MCP + Fine-tuning + Serving + Evaluation
FINAL BUILD+
Combine everything
- LLM
- RAG and vector database
- Tool calling
- Agents and LangGraph
- MCP
- Fine-tuned model
- Quantized inference
- API and Docker
- Evaluation and monitoring
- DGX Spark deployment/benchmarking
The flagship system should demonstrate that the roadmap became engineering capability — not just a list of technologies.
What every topic must earn.
The roadmap protects time by forcing every important concept toward practical capability.
Learn
Understand what the technology is, why it exists and where it fits in an AI system.
Build
Turn the concept into working code instead of collecting passive knowledge.
Debug
Understand common failures, bottlenecks, incorrect outputs and system behavior.
Explain
Be able to communicate the concept clearly in technical discussions and interviews.
Optimize
Measure memory, latency, throughput, quality and cost — then make tradeoffs.
Deploy
Move from notebook experiments to APIs, containers, GPU inference and production architecture.
Five projects that prove the skill.
The portfolio grows with the roadmap instead of becoming a separate task at the end.
Production RAG Assistant
Document upload, chunking, embeddings, vector search, local LLM, citations, evaluation and API.
Tool-Using AI Agent
A multi-step research agent that searches, reads, calculates, summarizes and answers using tools.
Multi-Agent + MCP System
Researcher, writer, reviewer and supervisor agents connected to custom MCP capabilities.
Fine-Tuned Domain LLM
Dataset preparation, LoRA/QLoRA training on DGX Spark, evaluation, quantization and inference.
Flagship Production AI System
A complete system combining RAG, agents, MCP, fine-tuning, serving, Docker, evaluation and monitoring.
Know when to stop.
The goal is engineering mastery, not research-level depth in every interesting topic.
Build + debug + explain + use. Examples: PyTorch, Transformers, LLMs, RAG, agents, MCP, LoRA/QLoRA, quantization, serving, Docker, evaluation and system design.
Understand the concept and use it when a project requires it. Examples: advanced CNN variations, RNN/LSTM, OpenCV, object detection, MLflow and cloud services.
Do not spend roadmap time here unless a project needs it: advanced proofs, research-level Transformer theory, reproducing papers and designing new architectures.