Projects

Project — Active

Local
LLM.

Self-hosted AI inference with llama.cpp (llama-server) on 3x NVIDIA Tesla V100 GPUs (96GB VRAM). One resident 35B MoE model at ~80 tokens/sec serving development, personal assistant, and content generation — zero cloud dependency.

llama.cpp / AI / Self-hosted / NVIDIA V100

$200-400
Monthly API savings
~80 tok/s
Inference speed (measured)
100%
Data stays local
4
Parallel AI sessions

01 — Configuration

The hardware.

GPUs 3x Tesla V100-PCIE-32GB
Total VRAM 96GB
Speed ~80 tok/s (measured on-box)
Access LAN only — not exposed

02 — Case study

GPU-accelerated local AI infrastructure.

The challenge

Run production AI inference for a personal assistant (Nova), 4 parallel AI coding workers, blog generation, and voice transcription — without recurring cloud API costs eating into a bootstrapped business budget.

The solution

  • Deployed 3x NVIDIA Tesla V100 GPUs with llama-server (llama.cpp) serving a 35B MoE model across all 3
  • Model chosen by measured head-to-head eval against larger dense models — the 35B MoE won on quality-per-watt
  • GPU-accelerated Whisper (large-v3-turbo) for voice transcription on dedicated GPU
  • Integrated with Claude Code workers — summaries and analysis patterns run locally for free
  • Nova personal assistant routes queries through 3 tiers: regex (free) → local LLM (free) → Claude API (paid)
  • Weekly blog posts auto-generated on the same resident model

03 — Available models

What's loaded in VRAM.

qwen3.6-35b

35B MoE (~3B active) / Always loaded across 3x V100 — no cold starts

The one resident brain — coding, reasoning, tool calling, Nova assistant, blog generation, classification

04 — Benefits

Why bother self-hosting.

Cost Reduction

Free inference for simple tasks that would otherwise use paid API calls. Saves money on summarization, status checks, and simple questions.

Privacy

Sensitive code and data never leaves the local network. No cloud provider sees your prompts or responses.

Speed

Local inference with no network latency. Responses start immediately without waiting for API round-trips.

Availability

Works offline and during API outages. Not dependent on external service availability.

05 — Integrations

Wired into everything.

Claude Code Workers Workers route summaries, classification, and analysis patterns to the local server before reaching for paid APIs
Nova Assistant Personal AI assistant uses the local server as its primary model with Claude API fallback
LLM MCP Server Model Context Protocol server for standardized local-LLM access
Blog Generation Weekly blog posts auto-generated on the local model via cron pipeline
Whisper Transcription GPU-accelerated speech-to-text for voice input via Nova

06 — Resources