Simon Schmincke 1 October 2026

A VC's AI Brain: Moving from Answering to Acting

The brain stopped waiting to be asked. Agents now run around the clock on a ten minute heartbeat, drafting my replies and repairing the system when it breaks. Nothing they write reaches anyone until I have approved it.

A follow-up to A VCs AI Brain, published in June. Same system, same chapters, three months further on.

Six building blocks that turn a VC's AI Brain from a search box into a working team: world model, agents, auto-reply, self-repair, apps and security.

01 | Since the last post

The reaction to the first post stunned me. It set off more conversation than anything else I have written: on calls, at conferences, over dinner, with some of the most forward-leaning, AI-pilled people I have met. Most of it was not praise, it was argument about what is actually technically possible today. Exactly the conversation I wanted. To everyone who reached out: thank you, and keep it coming.

So here is a proper update on what the thing has become since.

That June post described a retrieval system I had built for myself. It ingested every email, message, meeting, payment and document I touch, and answered questions about my life with full context. Two sentences in it are no longer true.

The first: "the core stack runs on my laptop and is reachable only from localhost."

The second, from the section on what the thing deliberately was not: "There is no orchestration agent. No meta-controller deciding which agent runs when. Today the orchestrator is me, sitting at the keyboard, deciding which conversation to open."

Both are now wrong. It runs on a machine of its own that never sleeps, I reach it from my phone, and there are agents on it doing work while I am not there. One of them finds what is broken overnight and fixes it. None of them can send anything to anyone until I have approved it on my phone.

That is the whole change. Everything below is either a piece of it, or a piece of the machinery that makes it safe enough to leave running. Same chapters as the last post, three months further on.

One thing I got badly wrong at the start, and it shapes everything below. I assumed the hard part would be making an agent capable. It takes an afternoon. What took the rest of the quarter was deciding what it may touch, proving that boundary holds when the model is wrong, and being able to read back afterwards what it actually did. That is where all the risk lives.

02 | The Machine

Everything is ingested, held and decided on one machine that never sleeps. The phone, the desk and the machine itself are viewports over it, and none of them holds data.

For the first year the whole thing ran on my laptop. Every operational problem I had was a version of the same sentence: the laptop was asleep. Close the lid on a flight and the jobs die. Travel for a week and the ingestion has a week-shaped hole in it. Ask whether the system is up and you are really asking whether I am at my desk.

Since August it runs on a Mac mini that never sleeps. Not a server in any impressive sense, just a small computer that is always on, and that turns out to be the entire feature.

Everything scheduled runs there: 107 active jobs, 47 databases, every service, every message bridge, every backup, all the monitoring and every send. My laptop is now a client. My editor opens a session that runs on the mini, same paths, same data, and I can close the lid mid-sentence.

Two rules came out of that move and most of the design follows from them:

One writer. One machine writes the databases. The laptop's copies are frozen at the migration date and every session is told not to read them. A stale copy looks exactly like the real one, and reading it and reporting the answer as current is invisible when it happens.

One exit. There is one way for a message to leave and it needs a tap on my phone. Not one way per channel. One way in total.

A written rule keeps it from decaying: everything new goes on the always-on machine by default, and the laptop keeps a short closed list of exceptions, work that needs a screen, models too large for the mini's memory, two helpers that only make sense next to the editor. A guard refuses to run a machine job on the laptop, so something installed in the wrong place fails loudly instead of running quietly in two places.

03 | The Network

Every service listens on the loopback interface, so the private mesh is the only door and identity arrives with each request. The open internet has no path in at all.

The June post said the stack was not exposed to the internet, and then put a "(yet)" after it. I want to take the "yet" back. How do you reach a machine from your phone without putting it on the internet at all?

Every service listens on the loopback interface only. Not the local network, not a forwarded router port, not a public address. If you were standing on my home network you could not reach any of it. There is no public endpoint anywhere, and that is not hardening applied afterwards, it is the decision that makes the rest affordable.

Reaching a loopback service from my phone is the job of a private mesh network. All my devices are joined to one, and it publishes a service as an encrypted endpoint on that mesh only. The feature that would expose it to the open internet exists and is deliberately never used.

Four things this buys.

Identity arrives with the request. The mesh signs each request with the identity of the device it came from, and every service checks it against mine. There is no password, no session cookie and no login form anywhere in this system, so there is no credential to phish, reuse, leak or rotate. The device is the credential and it was enrolled once, by me, on my own hardware.

A request with no identity is a local request. Because the sockets are loopback, anything arriving without a signed identity can only have come from a process already on the machine. The two cases separate cleanly.

Encryption without certificate management. Installable phone apps, background workers and push notifications all refuse to work over an unencrypted connection, and the mesh terminates a real one for me. Half of the phone chapter exists because that cost nothing.

My laptop reaches it identically. A tunnel forwards the same loopback ports, so code written against a local address behaves the same on either machine. Nothing in the codebase knows which host it is on.

What I gave up is real. Nothing reaches this system from a device that is not on the mesh, so if my phone drops off it the apps are dead rather than degraded. For a personal system that is the right trade.

04 | What it knows

One question reaches eight corpora through a single search. Results come back grouped per source, so half a million messages cannot bury a hand-written note.

The retrieval core is the part of the June post that changed least. I have not had to touch it in three months. One search over eight corpora: my notes, both mailboxes, WhatsApp, iMessage, calendar, the people graph and meeting notes. Around 640,000 indexed chunks. Keyword and vector search combined, results grouped per source so half a million messages cannot bury three hand-written sentences.

Two sources joined since June. Meeting notes, with attendees resolved against the people graph, so a bare first name in a transcript becomes the person I actually know. And my own promises, extracted deterministically from mail I have sent, in English and German, with hard exclusions for questions and conditionals. Twenty-eight things I said in writing that I would do, that I no longer have to remember saying.

05 | The World Model

Thirty-six kinds of fact, grouped. The kinds are the real vocabulary the system uses; the values on this card are illustrative and describe nobody. Every fact carries the verbatim sentence it came from.

This grew more than anything else in three months, and I would build it first if I started again.

It is not a CRM. A CRM holds the fields I bothered to type. This holds what the system has worked out about the people in my life from my own correspondence, and I type almost none of it.

There are 20,589 people in it, everyone I have exchanged mail or messages with plus everyone named in documents I hold. Around 1,300 of them carry extracted facts, which is the honest number: the fact layer covers the people who actually recur, not everyone who ever appeared in a thread. Between them they carry 5,730 facts in 36 kinds.

The kinds are the interesting part, because they are not the fields a contact database would have.

  • Tastes.

    What someone eats and drinks, the team they support, their sport, their pets.

  • Views.

    What they think about something, and what they have been reading. This is the largest group after our shared history, which surprised me.

  • Preferences.

    That one prefers a call to a thread. That another goes by a short form of their name and has done for twenty years.

  • Personal details.

    Children, partner, where they live, where they are from, which languages they speak, the dates that matter.

  • Working life.

    Current role, the role changes (the second largest kind, because people move), board seats, what they invest in, what they are expert in.

  • Our own history.

    How we met, what we have done together, what they have asked me for, and what I said I would do about it.

Three rules hold it together.

Every fact carries its evidence. The verbatim sentence it came from, capped at a couple of hundred characters, which channel said it, and when. A fact with no quote behind it does not get written. That is the difference between a system I can trust in front of someone and one that produces confident sentences I have to go and check.

Facts are superseded, never overwritten. Somebody changes job and the old role stays, marked as no longer current, with its own evidence and dates. The history is the useful part: knowing where someone was in 2021 is often the reason I can place them at all.

What I say outranks what it inferred. A fact I state myself carries full confidence and wins over anything the extractor worked out on its own. The system is allowed to be wrong; it is not allowed to argue with me about it.

Why does this earn a chapter of its own? Because almost every question about my working life reduces to a person.

Picture the two minutes before a call. You know this person is vegetarian, runs, has two small children, moved companies in March, thinks the seed market has gone mad, and asked you six weeks ago for an introduction you still have not made. That is worth more than any amount of document retrieval, and it is the difference between a warm conversation and a competent one. It used to mean reading six threads before dialling in. Now it is one card.

The change I would most defend in it is a subtraction. Fact extraction is the one step that posts my own message bodies to a model API, and confinement does nothing about that: confinement stops a model acting, and this is content leaving the machine. So that reader now answers a different question first. Not what the model may do with this, but whether it may be read at all. Two bodies of text are out of the world model permanently. The check happens on the same per-person lookup every read already goes through, so a button in the app cannot reach further than the nightly job can. And it refuses safely: if the exclusion list cannot be read, the whole run stops rather than falling back to reading everybody, because losing that list quietly would undo the whole point of having it. It records how many it skipped and never who, since writing the names down would put them straight back in a file.

06 | The Decision Layer: 24/7/365 on a ten minute heartbeat

A fixed loop reads my mail, messages and calendar, a drafter with no tools writes the text, and plain code decides whether it becomes a question for me. No step in it can send anything.

Every ten minutes. Around the clock, all year, whether I am at my desk, asleep or on a plane with no signal. That is 52,560 runs a year, and I see almost none of them.

This is the chapter that overturns the second quote at the top: the orchestrator is me, sitting at the keyboard.

The loop itself is deterministic and holds no model at all. It applies the decisions I have made, reads the signal caches its readers maintain, wakes anything snoozed, runs the proactive checks (stale threads, overdue promises, missing board material, calendar conflicts) and queues drafting work. Readers cover mail, messages, calendar and the promises I have made.

Then a drafter writes text. It has no tools, no network and no ability to send. Its entire output is words, handed back to trusted code, and that code decides whether anything becomes a question for me.

Four properties, none of which was obvious to me in advance.

Events exist even when nothing happens. Some events have no task attached. "The system saw this and decided to do nothing" is a record I want, and it is structurally invisible in any design where an event only exists as the justification for an action.

Every state change carries a mandatory reason. Not a log line, a required field. A card that cannot say why it surfaced does not get written.

Recipient provenance is checked, never inferred. A reply's recipients must appear in the source thread, verified against the mail store. A new message's recipients must come from a structured contact record. An address that exists only in the body of an incoming mail can never become a recipient, which closes the most obvious attack on a system that reads mail and can write it.

The one real playbook is mostly refusals. Introducing two people is three separate messages, each with its own approval. The first two may never name each other's recipient. The third is the only one where both appear, and it is unlocked solely by the target's own recorded agreement, read out of the mail corpus and re-checked at each stage. The quoted history is cut off before that agreement is read, so a reply quoting my own question can never be mistaken for consent to it.

It also learns where to stay quiet. A mail where I am on copy and the thread is informational gets a card and no draft, with the reason written on the card.

07 | The Agents

Every agent gets the same three tools. Which tables it may read, whether it may browse, and which actions it may run are one declaration per agent, so adding an agent is an entry in a file.

There are eight: mail, people, portfolio, restaurant, shopping, health, finance, and the caretaker in the next chapter. Portfolio is the newest, and the only one besides restaurant and health that can write anything.

For most of the quarter they were not really agents at all. Exactly one of them could write anything, through a branch hardcoded into the thread with its own action table. None of them could query a database, because a query needs a shell and a shell is the one thing these turns must never have. So the restaurant agent could read the entire internet and could not tell me when I was eating on Friday!

Only three of the eight can write anything even now: restaurant books, cancels and logs a visit, health logs food and training, portfolio writes a follow-up. The other five read and report.

They now share one chassis. Each gets the same three tools, scoped differently: one returns the table definitions it may read, one runs a read-only query over those tables, one runs a named action performed by plain code. The whole scope of an agent is a single declaration, which tables, whether it may browse the web, which actions it may run. Adding an agent is an entry in that file rather than a new code path.

Four decisions in it are the ones I would defend.

An agent's identity comes from where its connection lands, never from what the turn says about itself. Trusted code wrote that address before the turn started. An argument is a claim the model makes. Those are not the same kind of fact.

The boundary is the table, not the database it sits in. Granting at the database level would have been one line shorter and would have handed the mail agent everything that happens to live beside the mail.

Read-only is enforced twice, by the connection and by a check on the statement, because one of those will eventually be wrong.

Nothing here can send. No action may carry a send handler, and the module refuses to load if one appears.

Six of the eight may browse the web. The mail agent may not, deliberately, and neither does the caretaker. An agent that reads my mail must not also be able to fetch an arbitrary address, because that is the exact pair that turns a message somebody sent me into an instruction that leaves the machine. It is the cheapest way I know to close the biggest hole in a system like this, and it costs that agent almost nothing, since everything it needs is already on disk.

08 | Watching Itself

Detection is deterministic, execution is deterministic, and the one model turn sits between them writing a hypothesis that plain code validates. Seven topics can produce nothing but advice.

The monitoring was already good in June. Every job graded strictly, output freshness asserted rather than exit codes trusted, a push when something went red and another when it recovered. What it had no idea about was history. How long has this been broken? The alerts fired and then forgot, so the only honest answer was to go and read the logs.

There is a caretaker now. Every fifteen minutes it checks seven classes of fault, opens an incident for anything broken, appends an observation on every tick it is still broken, and closes it when the signal comes back clean. Those observations are append-only, enforced by the database, because they are the evidence for the duration and nothing that writes them should be able to tidy them up.

It re-derives nothing. The health grade comes from the same function that already pushes my phone. A second opinion on "broken" would be a second definition of broken, and the one that counts is the one that wakes me.

No model runs in detection. The single model turn in the subsystem sits between detection and action: it writes a hypothesis, names the evidence it read, and either proposes one action from the allowlist or escalates to me. It gets no tools. Trusted code gathers everything it may know and pastes it in, and a missing field says "unavailable" rather than being left out, because a model that cannot tell "there is no history here" from "I did not look" will invent one.

Confidence has to be a number, and a word like "high" becomes zero rather than being rounded up into certainty. Every piece of evidence it cites has to be a field that existed in what it was given. Every failure to parse lands on "escalate", which is a correct answer and not a fault.

It also repairs now, from a short list written in a file: run a job, reload its definition, restart one service, terminate an orphaned process from a fixed set, push me a notice. The target is checked against a denylist of paths first, on the path as written and on its resolved form. Credentials, the safety rails, anything outbound, deletion, the backup store and version history can produce nothing but advice, by construction.

Two rules in it I would carry anywhere. The model confirms a target, it never supplies one: every argument is read out of the registry for that incident's own subject and the model's string is only compared against it. And ran is not fixed: every action records a check fifteen minutes out and then a soak period, and until that holds the morning report says "tried, unverified" rather than "fixed".

The other half of watching itself is knowing which model each job runs on. Ask yourself that about your own automations and see whether you can answer it. Six of mine never said, so each ran on whatever the command-line tool's default happened to be that week, and that default moves every time the tool upgrades. My morning briefing changed model three times in two months and nobody chose any of it. Nothing broke, which is the problem: a silent substitution in something I read every day at seven is the kind of change I would notice only as a vague sense that the writing had got worse, with no way to test the hunch. There is a policy file now, keyed on the same job labels the usage reporting already uses, so the thing that sets the model and the thing that reports on it speak one vocabulary. It fails open by design, an absent pin means the old behaviour, with one inversion: the caretaker's diagnosis refuses to run at all if its pin is missing, because an unattended process that might later restart a service must never pick its own model.

09 | How Anything Leaves

The one path out, used by a background job at four in the morning exactly as by me at my desk. The approval covers one exact message, is spent when it is used, and no model sits anywhere in that path.

This is the part I would defend hardest and change least.

Every outbound message is blocked unless a matching approval already exists. Every email, internal and external, every message to another person, every post to a channel. The check sits at the tool-call layer rather than inside any one feature, so a background job at four in the morning and me at my desk both have to ask.

The approval is bound to the bytes. The token is a hash over the recipient, the body, the format and the sending identity. It authorises one specific message to one specific person, change a comma and it is void. It cannot drift into authorising something adjacent, because it never described an intention, only a message.

It is one-shot. The marker is consumed at the moment of sending. There is no standing permission anywhere in the system.

It binds the sender. I have two mail identities, and an approval for a private message must never be claimable by a work send with the same text.

No model sits in any send path. Not one. Once the deterministic checks pass, plain code consumes the approval and calls the mail API. This was forced by the same failure twice: a confined model refused a send I had already approved, calling it a suspected injection, and burned my authorisation doing it. Once I have tapped, the recipient and the bytes are fixed and all that remains is one call. There was nothing left for a model to decide, only something for it to get wrong.

Authorisation is not execution. An audit runs after the send and records what actually happened, and a watchdog raises an alarm if an approval is still unexecuted shortly after I gave it. Otherwise an approval gets spent by a send that then fails quietly.

Success is silent. A notification per delivered message confirms what I already know and trains me to ignore the channel that matters, so a delivered send just marks its card where I am already looking.

Attachments have their own refusal chain, because a file ships a whole document rather than a sentence. Identity documents are unattachable by any path for any recipient, permanently, and the check runs on the fully resolved path so a link cannot launder one. Every file is fingerprinted into the approval and re-verified before the approval is spent.

One rule here matters more than any of the cryptography. Being refused is not the same as being told to wait. The send is held, the session stays open, and it tries again the moment I approve. A process that has already exited cannot use my tap, and a session that tells me to approve and then finishes its turn has wasted one.

10 | The Surfaces

Real screenshots, blurred at capture time rather than afterwards, so no unblurred copy of these frames exists. Approvals sits empty most of the day, because the queue only fills when something wants to go out.

Three places I actually meet this system: my pocket, my desk, and a machine nobody sits at.

Five apps on my phone: approvals, system health, work, health logging and the cockpit. None of them came from an app store. Each is a web app served off the mini, added to the home screen from Safari, and reachable only inside the private mesh. Why five and not one with tabs? Because a phone gives you one badge per icon, and "three approvals waiting" and "one job failed" are not the same number.

Two pieces of that were harder than expected. Push runs with no third party in the path: the service generates its own signing keypair and holds the subscriptions, and since phones drop those after a week or two of disuse the apps re-subscribe on every open. And approving uses a key the server does not have: the phone generates a keypair, the machine keeps only the public half, so it cannot approve anything on its own behalf.

Each agent has a thread that says what it did and what it needs. The threads are assembled from what the agents already wrote, with the repetition collapsed out.

One thread per agent. Each agent has a standup card and a conversation: what it did, what it needs from me, and a box to write back. Nothing new is written for it, the threads are assembled from what the agents already wrote, with the repetition collapsed out. That collapsing is most of the work. A job log is not a conversation, and eight agents writing honest run summaries produce something unreadable until repeats fold into one line with a count and the runs that did nothing are dropped.

The cockpit became an application. It was 139 hand-written pages in June. It is now an app with 58 routes and one design system, because the 139 pages were never 139 things, they were about eight applications wearing a page each. A sweep loads every route at desktop and phone widths and fails on stub text, a table with headers and no rows, a console error, or a page whose title does not match the page it claims to be.

And the working session itself now runs on the mini, which means I can attach to it from my phone. Not a chat app, the actual session. It is the same program as my desk session on the same machine, so anything it tries to send still stops and asks, on the same phone I typed the instruction on.

11 | Learnings

A few things this round taught me that the first post could not have.

I thought the hard part would be getting an agent to do something useful. That took an afternoon. The rest of the quarter went on a duller question: what is this thing allowed to touch, and how do I know afterwards what it actually did? If you are about to build one of these, budget your time the other way round from how I did.

Never let the model tell you who it is. Each of my agents connects on its own address, and code picked that address before the conversation starts, so there is no box in which a turn could type "I am the health agent". Same when the caretaker wants to restart something: the name of the job comes out of my own registry, and the model's version of it is only checked against that. It confirms. It never supplies.

Write down what the system refused to do, not only what it did. My repair tier ran for a whole night and changed nothing, and every dashboard looked healthy. What gave it away was that it had also refused nothing. A thing that says no to everything and a thing nobody is asking look identical in a summary. If you keep the noes, they look nothing alike.

I had 1,544 of my own corrections and not one rule coming out of them, and I assumed the machinery was broken. The research says the opposite: rules a machine writes for itself make the thing measurably worse. So they were never rules. They are a test set, and now I replay them.

Run the same test twice before you believe any number from it. Mine moved by a hundredth with nothing changed in between. Without knowing that, I would have spent weeks celebrating improvements that were noise.

Keep the model in the middle and never at either end. Nothing it says decides whether something is broken, and nothing it says performs the fix. It writes a theory in between, and plain code checks the theory. Whenever that check fails the answer becomes "ask Simon", which was the right answer anyway.

The prediction I got most wrong was about local models. The June post gave them their own chapter and said my dependence on frontier APIs would go down, not up. Three months on, the local model layer is installed on neither machine and the job that used it is retired. The always-on machine has 16 GB and cannot hold a big model, the confined frontier calls turned out to be cheap enough, and the work simply moved. What stayed local is the embedding service, warm all day, serving every search. I would still bet on the original direction eventually. I was wrong about when, and about which half of it mattered.

12 | Next

The June post listed auto-memory curation and a procedural memory shelf. Both are still open, and they have now lost to infrastructure work two quarters running, which is data about my priorities rather than about their importance.

  • It notices, and it still mostly waits.

    The proactive checks run and the readers see things, and the system rarely opens a conversation about what it saw.

  • Repairs are not verified yet.

    Every one records when it should be checked and how long it must hold, and nothing fills that in, so the report honestly says "tried" rather than "fixed".

  • Calibration.

    The system records what I approve, edit and reject and does not yet use it to decide how confident to be.

  • Multiplayer.

    One person uses this. A team version needs per-user and per-record rights, a shared graph underneath and a personal layer on top, and none of that exists.

  • Native apps.

    I want to build my own mail client, on the phone and on the desk. The web apps have taken me a long way, but a native one gets proper notifications and the deeper system integrations that a browser tab is never going to have.

Closing

Same ask as last time. I am sharing this to inspire others, but really to learn from you if you have built something better. Reach out and teach me. Happy to buy you a coffee and geek out about how you have wired yours.

The thing I would most like pushed back on is what I say above about corrections. I think treating them as test cases rather than as raw material for rules is right, and it is a conclusion I reached from three papers and one month of my own data, which is not very much of either. If you have run that experiment properly, I would like to hear that I am wrong.

The June post said most of what was in it would be obsolete within months. It was: the two sentences I opened this one with went first. If you want to bet on which sentence from this post goes next, mine is "one person uses this".

Further
articles

Simon Schmincke
2026-06-22
A VC's AI Brain: Part One
A technical deep-dive into SimonOS: The tech stack that brings order to chaos, and handles the complexity that comes with a life lived across ten inboxes, two calendars, and 20,00 contacts.
Creandum Team
2024-09-26
Johan Brenner
2025-09-10
Reflections on Klarna: The Paper Invoice
From humble beginnings to European powerhouse
Creandum Team
2023-11-16
The Hottest Tech Ecosystem at Slush? (Hint, it's not Finland)
A transformation of the Lithuanian tech scene
Johan Brenner
2025-09-10
What would it take for the next Klarna to IPO in Europe?
Making Europe an attractive IPO destination