Skip to main content
AI / Machine Learning

Phase 5: Local AI & RAG Systems

10 Modules
Chapter 8: Resources: Local AI Models80%

Resources: Local AI Models

Comprehensive resource guide for mastering local AI.

Official Documentation

Ollama

LM Studio

llama.cpp

Model Sources

Hugging Face

Primary hub for downloading models:

Popular Model Cards:

Model-Specific Resources

Llama Family (Meta)

Mistral

Qwen (Alibaba)

DeepSeek

Hardware Guides

GPU Selection Guide

Entry Level ($200-400)

GPUVRAMBest ForPrice
NVIDIA RTX 30508 GB3B-7B models$250
NVIDIA RTX 306012 GB7B-13B models$300
AMD RX 6600 XT8 GB7B with ROCm$280

Mid-Range ($400-800)

GPUVRAMBest ForPrice
NVIDIA RTX 4060 Ti16 GB13B-30B models$550
NVIDIA RTX 407012 GB7B-13B (fast)$600
AMD RX 7900 XT20 GB30B models$750

High-End ($800-2000)

GPUVRAMBest ForPrice
NVIDIA RTX 408016 GB30B models$1,100
NVIDIA RTX 409024 GB70B models$1,600
AMD RX 7900 XTX24 GB70B models$1,000

Professional ($2000+)

GPUVRAMBest ForPrice
NVIDIA A600048 GB70B+ models$4,500
NVIDIA L40S48 GBProduction$7,500
NVIDIA H10080 GBResearch$30,000+

Apple Silicon Performance

M1/M2/M3 Family

ChipMemoryRecommended ModelsPerformance
M1 (8 GB)8 GB3B-7B Q412-15 tok/s
M1 Pro (16 GB)16 GB7B-13B Q418-22 tok/s
M1 Max (32 GB)32 GB13B-30B Q425-30 tok/s
M2 Pro (16 GB)16 GB7B-13B Q522-28 tok/s
M2 Max (64 GB)64 GB30B-70B Q430-40 tok/s
M3 Max (128 GB)128 GB70B+ Q445-55 tok/s
M4 Max (128 GB)128 GB70B+ Q450-60 tok/s
M4 Ultra (192 GB)192 GB70B+ Q855-70 tok/s

Resource: Does it ARM? - Check M-series compatibility

CPU-Only Systems

Minimum Specs:

  • 16 GB RAM: 3B-7B Q4
  • 32 GB RAM: 7B-13B Q4/Q5
  • 64 GB RAM: 13B-30B Q4
  • 128 GB RAM: 30B-70B Q4

CPU Recommendations:

  • Intel: i7-12700K or higher (12th gen+)
  • AMD: Ryzen 7 5800X or higher (Zen 3+)
  • Cores: 8+ cores recommended
  • Cache: Larger L3 cache improves performance

Budget Build (~$800):

CPU: AMD Ryzen 7 5700X - $180
Motherboard: B550 - $120
RAM: 64 GB DDR4-3600 - $180
Storage: 1TB NVMe - $80
Case + PSU: $150
Total: ~$710

Performance Build (~$2000):

CPU: AMD Ryzen 9 7950X - $550
Motherboard: X670 - $250
RAM: 128 GB DDR5-6000 - $450
GPU: RTX 4060 Ti 16GB - $550
Storage: 2TB NVMe - $150
Case + PSU + Cooling: $250
Total: ~$2,200

Learning Resources

Tutorials & Courses

Beginner:

Intermediate:

Advanced:

Video Tutorials

YouTube Channels:

Specific Videos:

Books

Free Resources:

Paid Books:

  • "Build a Large Language Model" - Sebastian Raschka ($40)
  • "Designing Machine Learning Systems" - Chip Huyen ($60)
  • "Natural Language Processing with Transformers" - O'Reilly ($50)

Tools & Libraries

Python Libraries

# Core
pip install ollama                 # Ollama Python client
pip install llama-cpp-python       # llama.cpp bindings
pip install transformers          # Hugging Face

# API Frameworks
pip install fastapi uvicorn       # REST API
pip install streamlit             # Quick UIs
pip install gradio                # ML interfaces

# Utilities
pip install pydantic              # Data validation
pip install python-dotenv         # Environment vars
pip install rich                  # Beautiful terminal output

Development Tools

Model Management:

Model Conversion:

Benchmarking:

Monitoring:

  • nvitop: GPU monitoring - pip install nvitop
  • htop: CPU monitoring
  • Prometheus: Metrics collection
  • Grafana: Dashboards

Communities

Forums & Discussion

Reddit:

Discord:

Forums:

GitHub Repositories

Essential:

Awesome Lists:

Benchmarks & Leaderboards

Model Performance

Official Benchmarks:

Hardware Benchmarks:

Comparison Tools

Papers & Research

Foundational

  1. "Attention Is All You Need" (Transformers)

  2. "Language Models are Few-Shot Learners" (GPT-3)

  3. "LLaMA: Open and Efficient Foundation Models"

Quantization

  1. "GGML - GPT-Generated Model Language"

  2. "QLoRA: Efficient Finetuning of Quantized LLMs"

  3. "AWQ: Activation-aware Weight Quantization"

Optimization

  1. "FlashAttention: Fast and Memory-Efficient Attention"

  2. "Mixture of Experts"

Troubleshooting Resources

Common Issues

Issue: Out of Memory

Issue: Slow Inference

Issue: GPU Not Detected

FAQ Collections

Advanced Topics

Fine-Tuning

Tools:

Guides:

Model Merging

Tools:

Guides:

Deployment

Production:

Edge:

Datasets

Pre-training Data

Fine-tuning Datasets

Evaluation Datasets

Security & Privacy

Best Practices

Compliance

Stay Updated

News & Updates

Newsletters:

Blogs:

Twitter/X:

  • @ollama_ai
  • @ggerganov (llama.cpp creator)
  • @reach_vb (AI news)
  • @simonw (Simon Willison)

Conferences & Events

Contribute

Help improve these resources:

  1. Report Issues: Found a broken link or outdated info? Let us know
  2. Share Resources: Know a great tutorial? Submit it
  3. Community: Join discussions and help others
  4. Code: Contribute to open-source projects

Navigation

Quick Reference

Start Here:

  1. Install Ollama: https://ollama.com
  2. Join r/LocalLLaMA: https://reddit.com/r/LocalLLaMA
  3. Read "Illustrated Transformer": https://jalammar.github.io/illustrated-transformer/

Need Help:

Keep Learning:

  • Subscribe to newsletters
  • Follow key developers
  • Join community projects
  • Experiment and share findings

Remember: The local AI community is collaborative and welcoming. Don't hesitate to ask questions and share your experiences!