My AI drafting tool became a system the first time it opened a pull request without me.
Until then it was a utility. A script watched a handful of feeds, scored events, and handed me a prompt. I pasted that prompt into a chat window, edited the output, and published. Every decision that mattered happened in my head, and the AI never held a privilege I had not personally exercised that minute.
When I decided that Swiss Security Insights should run as an agentic newsroom, with five AI journalists, an AI editor and one human who approves everything, I stopped designing a prompt and started designing an organization. An organization has identities, scopes, credentials, untrusted inputs, a budget and a point of accountability. I have spent more than twenty years asking those questions about other people's systems. This time the system was mine, small enough that I could afford to get things wrong, and instrumented enough to learn from every mistake.
This article is the case study. It covers what I built, the decisions I made and why, what broke, and the framework I extracted from the experience. I report the numbers I have, including the unflattering ones. A series of follow-ups will test the system over the coming weeks.
I write this as a CISO, not as a prompt engineer. Whether a model can write an article is no longer an interesting question. The useful question is what you must build around the model so that you would sign your name under whatever comes out.
Phase 1: a tool that was safe because it was small
The first version was safe for a boring reason: it could not do anything.
A Python agent ran on a Mac mini under launchd. It polled eight feeds (CISA KEV, NVD, NCSC, ENISA and four trade and research sources), asked a model to triage each event, and scored its relevance from 1 to 10. Anything at 7 or above produced a ready-made prompt. I carried that prompt by hand into a chat session, edited what came back, turned it into an article page, and published through CI. The agent never published, never held a credential that could change the site, and never saw its own output again.
From a risk perspective that is a clean design. The model has no write path, the human is the only actor with privilege, and the blast radius of a bad output is one draft that I may or may not read carefully. The control is me.
That is also exactly why it stopped scaling. Three limits showed up together. The control did not scale: one person reviewing everything is a control until the volume grows, and after that it is a bottleneck that people quietly route around. The infrastructure was a single point of failure: if the Mac mini slept, the pipeline stopped, and nothing recorded what it had missed. And nothing was auditable: the prompt, the output and my edits lived in a chat window, so I could not show anyone, including myself a month later, why an article said what it said.
The tool was an accelerator for one person. To become anything more, it needed a trail, a perimeter and a division of labour.
Phase 2: six decisions that shaped the architecture
I made six architectural decisions before I wrote a line of agent logic. Each one traded something away, and I want to be explicit about what.
| Decision | What I rejected | Why | What it costs |
|---|---|---|---|
| Agents run on GitHub Actions on a schedule | The Mac mini, or a rented VPS | No machine of mine sits in the execution path, and every run starts from a clean runner with a full log | Dependence on a platform, and egress from shared runners that some sources block |
| Agents live in a private repository, the site in another | One repository for everything | The agent code and the production site no longer share a trust boundary | Two repositories to keep coherent |
| No long-lived model API key in CI | A static key stored as a secret | Workload identity federation means there is no standing credential to steal or leak | More setup, and a few days of debugging unfamiliar failure modes |
| The bot reaches the site only through a pull request | A personal token or a direct push | A GitHub App with minimal permissions, short-lived tokens and a protected main branch cannot publish on its own | Every article waits for a human |
| The pipeline is code with tests | Low-code automation, or an assistant-based workflow | Behaviour can be reviewed, versioned and tested, and there is no vendor lock-in on the logic | I maintain it |
| Start at one article a day with approval on everything | Higher volume with sampled review | Autonomy should be earned from evidence, not granted up front | Low throughput by design |
The fourth row deserves a closer look, because it is the decision I would defend hardest in front of a board. The bot is a GitHub App installed on a single repository with only the permissions it needs to propose a change. I enabled a ruleset on the main branch that requires a pull request and an approval, and then I tried to push directly. The push was refused. I did not trust the configuration until I had watched it say no.
I also changed how the model authenticates. In CI there is no long-lived API key. The workflow obtains a short-lived identity token from GitHub and exchanges it for access. It was the most instructive piece of the build: an early version created a fresh client for every call, which reused the same one-time token and was rejected as a replay. The control worked. My code was the thing that was wrong.
One more disclosure belongs here. I built this system with an AI assistant, and I treated that as a governance question too. I worked one problem at a time, tested each change on an isolated copy with a clean environment, ran every command myself, and committed only after the full test suite and linter passed. An AI that helps you build an AI system needs the same boundaries as the system itself.
Phase 3: a newsroom with a chain of custody
The newsroom has six AI roles and one human. Five journalists each own a domain, and an editor decides what is worth writing about.
| Role | Domain |
|---|---|
| Xander Ploit | Vulnerabilities and actively exploited threats |
| Regina Tory | Swiss and EU regulation |
| Brett Each | Incidents and ransomware |
| Ada Prompt | Security and governance of AI systems |
| Faye Lover | Operational resilience for SMEs |
| Phil Ter (editor) | Selection, scope enforcement and ranking |
The names are deliberate puns, and they are also a disclosure. Nobody should mistake these roles for people, and the byline of every article says that it was written by an AI and reviewed and approved by me.
The scopes follow domains, not the five categories of the site. That choice matters more than it looks. A journalist with a narrow, written beat can be told to skip a story that does not fit, and the editor can disqualify a pitch on scope before judging its quality. Both rules exist because the early versions of the system did the opposite: given a story and a persona, a model will find a way to write something. I wanted the default behaviour to be refusal.
The flow runs in one direction, and each stage can say no. Events from fourteen sources are pooled and clustered, so a vulnerability reported by three outlets becomes one story. A deterministic check against published and in-flight articles removes what SSI already covered. Each journalist may pitch at most one story a day. The editor ranks the viable candidates and, if the top one fails to write or fails verification, falls back to the next. Only a draft that survives verification becomes a pull request, and only I can merge it.
The framework: six boundaries of an agentic system
Every control I added falls into one of six boundaries. If you are about to put an agent anywhere near production, these are the six questions I would ask before the first run.
| Boundary | Risk it addresses | Control | Evidence from the first runs |
|---|---|---|---|
| 1. Identity and scope | Scope drift and diffuse accountability | Named roles with written beats, skip over force, editor disqualifies out-of-scope pitches | 155 pitch attempts across four runs produced 4 pitches; the rest ended in a logged skip |
| 2. Untrusted input | Feed content steering the model, the page or the CI log | Feed text treated as data, model output escaped before rendering, untrusted text confined to one log line, outbound fetches guarded against SSRF | Closed in the review (18 of 18 findings); the log-injection fix came from reading the second run |
| 3. Verification | Invented sources, figures and unsupported claims | Deterministic checks first (URLs, dates, durations), then an LLM claim check against fetched source text; a single source only if primary | Run 1: draft blocked with 7 findings, and the LLM check flagged the same claim independently |
| 4. Privilege and blast radius | An agent or a dependency reaching production | Separate repositories, minimal GitHub App, short-lived tokens, no standing model key, actions pinned by hash | A direct push to main was refused when tested |
| 5. Human accountability | Rubber-stamping and unclear ownership | Approval required by ruleset, AI byline with review statement, audit artifact for every run | Two bot pull requests reviewed and merged within about 10 minutes of each run |
| 6. Cost and failure containment | Runaway spend, one failure aborting a run, stories lost without review | Persistent daily and monthly budget, refusal of unpriced models, per-pitch failure isolation, capped attempts with carry-over | 7 of 154 development calls hit the length limit, which led to a fix in the writer and in its retries |
Three notes on how to use this. First, the boundaries are independent. A strong verification layer does not compensate for a weak privilege boundary, because they fail in different ways: one protects the content, the other protects the platform.
Second, they map onto frameworks a risk committee already knows. Boundaries 1 and 5 are mostly governance. Boundaries 2, 4 and 6 are about managing risk in operation. Boundary 3 is measurement. That is a rough mapping to the functions of the NIST AI Risk Management Framework, and it gives you a way to translate the exercise for a board without inventing a new vocabulary.
Third, the three lines model applies. The journalists are the first line, the editor and the deterministic checks the second, and I am the third. What is missing is independent assurance. I have not had this system audited by anyone but myself, and I would not describe it as assured until someone else has.
Nothing here is specific to media. A bank or an insurer deploying agents faces the same six boundaries, and the controls differ in weight, not in kind. Swiss supervisors have started to say so explicitly. FINMA Guidance 08/2024 is not binding, but it sets out what supervised institutions are expected to bring to AI: central accountability, independent review by skilled people, adversarial testing and monitoring, an understanding of how the application behaves, and proper management of third-party providers.
My newsroom does not satisfy that guidance, and it was never meant to. It is a deliberately small system, and it shows how little of the guidance is about the model and how much of it is about the structure around it. A later piece in this series will go through it line by line, together with the obligations that DORA and the EU AI Act add for institutions with EU exposure.
What went wrong, and what the numbers say
The first real run produced no article, and that was the system working.
The writer had delivered a clean draft of about a thousand words. It read well. It also cited a source whose URL was a placeholder the model had left instead of a real link, and it repeated an "eight-week" consultation window that nothing in the sources supported. The deterministic checker blocked it with seven findings. The separate LLM claim check flagged the same claim on its own. If I had been reviewing by reading, I might have missed both. This is the reason the first verification layer is deterministic: a fluent draft is the hardest kind to catch by eye.
The second run showed the cost of catching the problem late. A different journalist's draft was blocked for the same placeholder source, after the tokens for the full article had already been spent. I moved that check to the moment of the pitch, so a pitch with an unverified source is discarded before the writer sees it and the journalist moves on to the next story.
Three other problems are worth stating plainly. Two sources returned 403 errors to the hosted runners, so I added a single retry and a "skipped" status so that one blocked feed cannot fail the run, and I disabled one source that kept refusing; a pipeline that depends on public feeds inherits their anti-bot rules. Before the ingest index was persisted between runs, every scheduled run treated the same events as new, and a stateless job is easy to build and wrong in a way that only repeated runs reveal. And in development, 7 of 154 model calls stopped at the length limit, so the writer now has room for a full article and a retry revises the rejected draft instead of starting over.
Two problems remain open. The first is efficiency. Journalists spend most of their attempts on stories that are clearly outside their beat, and a hand-off to a colleague is lost when that colleague has already used their attempts. The second is that claim checks still raise warnings on interpretive sentences in the articles that did pass. Those warnings concerned analysis rather than figures, but a sentence that is interpretation and not fact is exactly where a human editor has to stay involved.
The system produced 2 articles from 4 pipeline runs, and rejected most of what it considered. These are early numbers from a small sample, and I am publishing them as observations, not as performance. The four runs below took place on the same evening of 1 October, while I was hardening the pipeline. They are not a steady state, and the series will report on one.
| Run 1 | Run 2 | Run 3 | Run 4 | |
|---|---|---|---|---|
| Duration | 3 min 41 s | 3 min 45 s | 2 min 44 s | 3 min 52 s |
| New events | 21 | 14 | 0 | 0 |
| Events in pool / clusters | 124 / 112 | 138 / 126 | 138 / 126 | 138 / 126 |
| Pitch attempts | 35 | 40 | 40 | 40 |
| Pitches reaching the writer | 1 | 2 | 0 | 1 |
| Drafts blocked by verification | 1 | 1 | 0 | 0 |
| Articles rendered | 0 | 1 | 0 | 1 |
Across the four runs, 155 pitch attempts led to 4 pitches that reached the writer, about 2.6 percent. Nearly every rejection carries a written reason: outside the journalist's beat, already covered by SSI, too thin to assess, or sourced from a placeholder. I read that rate as a control working, and also as an inefficiency I still have to fix. Run 3 ended with no story at all because there were no new events, and I consider "nothing worth publishing today" a valid outcome for a newsroom.
The rest of the picture is simple. The build took 71 commits in seven days, with 29 modules and about 3,900 lines of code, backed by about 3,900 lines of tests. My security and reliability review closed 18 of 18 findings, and the real runs surfaced ten more that I fixed. Fourteen sources are active, one of which blocks the hosted runners intermittently. Three of the 103 articles on the site were produced by the newsroom. Each of the two pull requests from these runs was merged about ten minutes after its run finished, an upper bound that includes the time for the pull request to open.
What I kept human, and how autonomy gets earned
I kept four things human on purpose, and I do not plan to automate any of them soon.
The first is the merge. Nothing reaches the site without my approval, and the platform enforces that, not a policy document. The second is judgement about interpretation. Models are good at stating what a source says and unreliable at deciding what it means for a Swiss bank or insurer, which is where the claim-check warnings keep landing. The third is scope: what SSI covers, and which of the five beats deserves a journalist at all. The fourth is the signature. My name is on the review statement, so I carry the accountability, and I read accordingly.
Loosening any of this has to follow evidence. I will not raise volume or reduce review on a feeling that the system seems reliable. Before any change I intend to look at four measures. Defect escape counts how many published articles needed a substantive correction after merge. Intervention rate tracks how often I change a draft, and how much, and whether that rate is falling. Catch rate shows whether the verification layers keep stopping the same classes of defect, which tells me they still work. Anomalies covers any security finding, unexpected credential use, budget breach or run that behaved outside its design.
If those hold over several weeks, the first relaxation would be narrow: sampled review for the lowest-risk category, with a rollback condition written down in advance. I commit publicly to reporting each of the four measures in this series as the data accumulates, including the weeks in which they get worse.
◆ Key Takeaway
Autonomy is a permission you extend to a system after it has shown you what it does, and it needs a mechanism for taking it back.
- Design the refusal path before the happy path. An agent that cannot say "not today" will produce something, and what it produces will be fluent. Make skipping the default and make every skip leave a reason.
- Put the deterministic control in front of the probabilistic one. A script that checks a URL, a date or a duration is cheaper, faster and more predictable than a second model, and it catches the defects a model is most likely to share with the first one.
- Treat both ends as untrusted. The input can steer the model, and the output can carry content into a page, a log or a pipeline. Apply the same discipline to both.
- Make accountability structural, then test it. A policy that says a human approves everything is a sentence. A ruleset that refuses a direct push is a control. Try to break it before you believe it.
- Measure from the first run. Without the logs of the first four runs I could not have told you where the system wasted effort, which defects it caught, or how long my review took.
- Write down the conditions for loosening review before you start. Decide in advance which evidence extends autonomy and which event takes it back.
- Get independent assurance before you call the system assured. The person who built the controls is the wrong person to certify them.
This is the first of a short series. The newsroom keeps running, and I will report what it does, including what I have to revert. The next pieces cover the threat model of a newsroom of agents, starting with prompt injection through feeds; why verification starts with deterministic checks and what the model-based layer adds; least privilege for agents in CI/CD, from tokens to supply chain; the governance of cost and failure in an agent system; what FINMA, DORA and the EU AI Act ask of a system like this; and a summary after several weeks of data, with the four autonomy measures and the decision they support.
The question I started with was whether an AI could help me write faster. The question I ended up answering is what it takes to be accountable for a system that writes under my name. The second one is the one that transfers to every organisation I would want to work with.