Локальные LLM на своём железе: через что запускать, какую модель выбрать и где польза

Vladimir Novator

Vladimir Novator

2,995 views

How to run your own ChatGPT directly on your GPU: no internet, no subscriptions, and no token limits. I'm analyzing local LLMs on an RTX 5090 and, in the end, assembling a swarm of three AI agents that write a game with a single prompt.

In the video:
Why run neural networks on-premises and when it doesn't pay off
Which models to choose (Qwen, Gemma) and the licensing pitfalls
How much video memory do you need: 8, 16, 24, and 32 GB
Quantization in simple terms: Q4, Q6, Q8
LM Studio, Ollama, or llama.cpp: which to choose
Why it's important to limit context
Agent swarm in OpenCode: orchestrator + reviewer + QA
Honest summary: where's the benefit and where's just a toy?

My Telegram channel: https://t.me/noovator

0:00 Introduction
0:18 What does it mean to run a neural network on-premises
1:17 Why is it necessary: ​​5 reasons
2:46 What can you run? Locally
3:35 Which models to take and the licensing pitfalls
4:19 What hardware do you need
5:26 Quantization in simple terms
6:22 What to run: LM Studio, Ollama, llama.cpp
7:36 LM Studio: overview and settings
9:11 Why context should be limited
9:33 Max Concurrent Predictions
10:09 Agent swarm: how it works
11:09 Demo: Tower Defense with one prompt
12:23 What the reviewer and QA found
13:45 Testing the game
14:57 Result: a game in 10 minutes
15:09 Models without failures
15:45 Honest result: where is the benefit and where is the toy
16:31 Where to start: 4 steps