top of page

DAKOTA INTELLIGENCE | Local-First AI: Taking Back Your Machine

  • Writer: Timm Johnson
    Timm Johnson
  • Apr 24
  • 10 min read

DAKOTA INTELLIGENCE

Automate South Dakota

LOCAL-FIRST AI

Taking Back Your Machine



A Plain-Language Guide & Technical Proof of Concept

Timm Johnson  |  Dakota Intelligence

Mitchell, South Dakota  |  2026



[ PART ONE — PLAIN LANGUAGE ]

The Felt Problem

Your computer is supposed to work for you.


That used to be obvious. You bought the machine, you owned what ran on it, and what happened inside that box stayed inside that box. Somewhere along the way, quietly, that stopped being true.

It didn't happen all at once. It was a setting here, an update there, a new feature nobody asked for. Screenshots taken automatically in the background. An account you didn't want becoming one you couldn't skip. Data leaving your machine for reasons described in a terms-of-service document no one reads. Small things, individually defensible. Together, they add up to a machine that works with you some of the time, and on you the rest.

For most people, that tradeoff is invisible. For some of us, it became impossible to ignore.

When you're handling client information — real people, real trust, real consequences if something leaks — you start looking at your own desktop differently. When you've spent years developing ideas you consider genuinely original, ideas you've deliberately kept off cloud infrastructure until you're ready, the question of where does this data actually go stops being abstract. When your work is built on the premise of helping people trust technology, you can't be the person who never looked too closely at what their own machine was doing.

That's where this started. Not ideology. Necessity.


The M365 Acknowledgment

Let's be clear about something before we go further.


Microsoft 365 is genuinely good. Word, Excel, Teams, Outlook — these tools work, they're familiar, and for most of the businesses and people you work with every day, they're the water they swim in. Telling someone to abandon M365 is like telling a carpenter to stop using the hammer everyone else on the job site uses. Technically possible. Practically a different conversation.

This isn't that conversation.

The argument here isn't that Microsoft is the enemy. The argument is simpler and less dramatic: a leaky container is still useful. You just don't store everything in it.

You know which drawer in your kitchen has the loose bottom. You still use that drawer — just not for the heavy things you don't want to lose. Same logic applies here. M365 lives on the machine. It does its job. But your client conversations, your original research, your AI inference, the thinking you haven't shared with anyone yet — that doesn't run through it. That runs somewhere you control completely.

The line isn't Microsoft versus freedom. The line is: what do I actually need to protect, and where does it actually live?

Once you ask that question honestly, the build starts making itself.


What You Built and Why

At some point, thinking about the problem stops being enough. You build something.


The machine sitting on the desk isn't exotic. It's a tower — a mid-range workstation built deliberately, not expensively. A modern processor, a capable graphics card, a fast drive. Nothing a small business owner or a serious hobbyist couldn't put together. The point was never the hardware. The point was what runs on it and what doesn't leave it.

The AI stack is local. That means the models — the actual intelligence doing the reasoning, summarizing, drafting, and inferring — run on that machine, on that graphics card, without a single token of your input traveling to a data center you don't own. No API call to someone else's server. No conversation logged somewhere in the cloud. You type, it thinks, you get an answer. Start to finish, your building.

Windows is still there. It has to be — clients use it, certain tools require it, the world runs on it. But it's been demoted. It's the surface layer, not the foundation. The work that matters — the sensitive client context, the original research, the AI reasoning over your own documents — that happens in an environment you've deliberately separated from the defaults.

It didn't require a computer science degree. It required being annoyed enough, long enough, to start asking different questions. What if the AI ran here instead of there? What if my notes never touched a server I don't control? What if I treated my own machine like it actually belonged to me?

Those questions have answers. Practical ones. That's what this build is.


What It Costs, What It Gives Back

Nothing worth building is free. That's not a warning — it's just honest.


Here's what local-first costs you. Setup time — not weeks, but real hours of learning, configuring, and occasionally breaking things on purpose to understand how they work. Some friction when a tool you're used to doesn't behave exactly like it did before. A periodic reminder that you are now the person responsible for your own stack. There's no help desk. There's no support ticket. When something doesn't work, you figure it out — which, depending on who you are, is either the worst part or the whole point.

Some things you took for granted get slower before they get better. That's real and worth saying.

Here's what you get back.

Your AI doesn't know who your clients are unless you tell it — and when you tell it, that information stays on your machine. Your research drafts, your half-formed ideas, your notes at midnight — none of it is training someone else's model or sitting in a retention policy you didn't write. When you close the lid, the work is yours. Completely.

There's also something harder to quantify but impossible to ignore once you feel it: the work gets quieter. Fewer background processes quietly phoning home. Fewer nudges toward features you didn't ask for. Just you, your machine, and the problem in front of you.

For a consultant, for a researcher, for anyone whose value lives in what they know and what they're building — that quietness isn't a luxury. It's infrastructure.


[ BRIDGE ]

Figure It Out Yourself

Nobody gives you the blueprint for the thing that doesn't exist yet.


There's a moment — most builders know it, even if they've never named it — where you're standing at the edge of a problem and the usual options have run out. You've searched. You've asked. You've waited for someone more qualified to show up and solve it. They didn't. The problem is still there, and now it's just you and it, and the uncomfortable silence between.

That moment is where everything real gets built.

It doesn't matter if you're standing in a field trying to figure out why the equipment failed three miles from anyone who can help, or sitting at a desk at midnight staring at a system that almost does what you need. The feeling is the same. No manual covers this exact situation. No forum thread ends with your answer. The expertise you need doesn't exist yet because the problem you're solving is specific enough, or new enough, or strange enough that you're the first person to need it solved in exactly this way.

Most people back up. Find a workaround. Accept the limitation. That's reasonable. There's no shame in it.

But some people get curious instead of defeated. They start pulling threads. They break the thing deliberately just to see what's inside. They build something rough that works, then build it again better. They make mistakes that cost them hours and learn things no tutorial could have taught them. They arrive at a solution that is entirely their own — not because they're exceptional, but because they stayed in the room long enough.

The innovation wasn't the destination. The stubbornness was.

That's true whether you're rewiring a barn, writing a theorem, or building an AI stack on a tower in South Dakota because the defaults stopped working for you and you decided to do something about it.

The blueprint comes after. First you just have to start.


[ PART TWO — TECHNICAL PROOF OF CONCEPT ]

Threat Model First

A build without a threat model is just vibes with a power cord.


Before anything else — before the OS choice, before the inference stack, before any of the decisions that follow — you have to answer one honest question: what are you actually defending against?

Not in a paranoid sense. In an engineering sense. A threat model is just a clear-eyed list of what you have, what could go wrong, and which risks you're willing to accept. Without it, every security decision is arbitrary and every tradeoff is invisible.

What's Worth Protecting

  • Active client context — names, situations, operational details shared in trust

  • Pre-publication research — original frameworks and ideas kept deliberately off cloud infrastructure

  • Business intelligence — leads, strategy, pricing, positioning

  • Inference conversations — prompts and outputs generated while working, which reveal as much as the underlying data

Realistic Threats

  • OS-level telemetry that vacuums behavioral data by default

  • Cloud AI providers with opaque or unfavorable retention and training policies

  • Accidental exposure through sync services that grab more than intended

  • Accumulation of behavioral data in places you don't control and can't audit

What This Build Accepts

Windows still exists on this machine. That's a known, managed risk — not an ignored one. The strategy is containment, not elimination. Sensitive workflows run in isolated environments. The attack surface is reduced deliberately and documented honestly.

What This Build Does Not Defend Against

Physical access. A compromised local network. Human error. This isn't a bunker — it's a disciplined workspace. Know the difference.


The Stack

Local-first doesn't mean primitive. It means intentional.


The OS Layer

Windows remains as the host environment — a deliberate, pragmatic choice. Client tools, productivity software, and hardware compatibility all require it. But the sensitive work runs inside isolated environments: WSL2 (Windows Subsystem for Linux) for development and AI workloads, with careful attention to what crosses the boundary. For anyone starting fresh or building a dedicated AI workstation, a Linux base — Fedora for familiarity, NixOS for reproducibility, Bazzite for gaming and GPU workloads — removes the host-level telemetry problem entirely.

The OS is a container, not a trusted partner.

The Inference Layer

Ollama is the engine. It runs open-weight models — Llama, Mistral, Phi, Qwen, and others — entirely on local hardware, using the GPU for acceleration. No API call leaves the machine. No conversation is logged externally. The models themselves are weights stored on your own drive, loaded into your own VRAM, producing output that never touches a network unless you explicitly route it there.

The Data Layer

Documents, notes, research, and client context that need to be reasoned over stay in a local vector store — a small database that lives on your drive and lets the AI search your own knowledge without any of it leaving the machine. Combined with local inference, this means you can ask questions about your own documents, generate summaries, draft from your own notes, and build context-aware workflows entirely within your own walls.


Build Decisions and Why

Every choice in a build is an argument. Here are the ones this build is making.


GPU Over CPU for Inference

The RTX 5060 in this machine wasn't chosen for gaming — it was chosen because modern AI inference is GPU-bound. Running models in VRAM is dramatically faster than running them in system RAM or on the CPU alone. An 8GB card handles most 7B and some 13B parameter models comfortably. That covers the majority of practical local AI workloads: summarization, drafting, coding assistance, document Q&A, and conversational reasoning.

Ollama Over llama.cpp Directly

Both work. Ollama wraps llama.cpp with a cleaner interface, model management, and a local API endpoint that other tools can call. That matters because the goal isn't just a chatbot — it's a stack. Make.com scenarios, Base44 apps, and other automation layers can call a local Ollama endpoint the same way they'd call a cloud API, with the critical difference that the traffic never leaves the machine.

WSL2 as the Isolation Boundary

On a Windows host, WSL2 provides a Linux environment that runs the AI stack cleanly, with good GPU passthrough support via CUDA. It's not perfect isolation — Windows can still see the filesystem if you're not careful — but it separates the inference environment from the Windows application layer meaningfully. File storage for sensitive workloads lives inside the WSL2 filesystem, not in Windows-accessible paths that sync tools might grab.

Local Vector Store Over Cloud RAG

ChromaDB or a similar lightweight vector database runs inside the local environment. Documents get chunked and embedded once, stored locally, and retrieved locally. The embedding model runs on the same GPU. The entire retrieval-augmented generation pipeline is self-contained. This is the piece that makes local AI genuinely useful for knowledge work — not just a chatbot, but a reasoning layer over your own accumulated context.

No Mandatory Microsoft Account

Where possible, local accounts only. Telemetry settings pushed to minimum. No automatic submission of diagnostic data. This doesn't eliminate Windows telemetry entirely — on Home and Pro that's not fully possible — but it reduces the signal substantially and documents the decision honestly.


What Windows Still Touches and How You Contain It

Containment is not the same as elimination. Know the difference and work with it.


Windows still sees the hardware. It still handles drivers, networking, and the application layer for productivity tools. That's the surface you accept when you keep Windows as the host. The strategy isn't to pretend otherwise — it's to be deliberate about what runs where.

Sensitive AI workloads run in WSL2, not in Windows applications. Notes and research live in local directories inside the Linux environment, not in OneDrive-synced folders. Browser profiles used for client work are separated from general browsing. Cloud sync tools — OneDrive, Google Drive, Dropbox — are either disabled or scoped explicitly to non-sensitive directories.

The mental model is zones. Windows gets the outer zone. The inner zone — inference, research, sensitive client context — runs local-first by design, and the boundary between zones is maintained consciously.

It's not perfect. It's honest. And it's substantially better than the default.


Reproducibility Notes

If you want to build this, here's where to start.


Hardware Minimum

  • Dedicated GPU with 8GB VRAM minimum (handles most 7B–13B parameter models)

  • 16GB system RAM

  • NVMe drive for model storage and fast load times

  • Mid-range consumer hardware — nothing exotic required

Software Stack

  • Ollama (ollama.com) — model management and inference engine

  • WSL2 with Ubuntu or Fedora — Linux environment on Windows hosts

  • ChromaDB or Qdrant — local vector storage

  • Open-weight models via Ollama: Llama 3 or Mistral for general work, Phi-3 for lighter tasks

The Discipline

Tools are the easy part. The discipline is the practice — knowing which work goes inside the walls and which doesn't, maintaining that boundary when convenience pushes against it, and periodically auditing what's actually leaving your machine. Set up a local network monitor once. Watch the traffic. You'll stop being surprised quickly, and you'll make better decisions after.

The Payoff

A machine that reasons over your documents without reading them to anyone else. An AI assistant that doesn't know your clients' names unless you're the one who told it. A workspace where the thinking you haven't published yet stays that way until you decide otherwise.

That's not a luxury setup. For anyone whose work lives in what they know and what they're building — it's just the right tool for the job.



DAKOTA INTELLIGENCE

Automate South Dakota




Recent Posts

See All
bottom of page