Closing the loop: a self-improving AI bot for network access automation
There is a category of engineering work that no one is proud to own: the repetitive, low-ambiguity, high-volume operational task. Network access requests are a perfect example. A ticket arrives describing what a user or system needs to reach. Someone reads it, figures out what IP ranges are involved, translates that into a firewall rule, opens a review, watches the CI pipeline, merges it, and moves on — only for the next one to arrive an hour later.
It is not hard work. It is just constant work. And it is exactly the kind of work that AI coding agents should own.
This article describes how we built a bot that handles the complete lifecycle of network access requests — from reading the ticket to watching the deployment pipeline — and, crucially, how it improves its own behavior after every failure without human involvement.
The problem: structured repetition at scale
Network access automation looks simple from the outside. The inputs are constrained: a request says who needs access, to what service, over which protocol. The output is constrained: a rule in a configuration file, reviewed through GitOps, deployed through CI. The feedback signal is unambiguous: the pipeline either passes or it doesn't.
But in practice, several things make this harder than it looks:
- Variable resolution: The ticket names a service. The firewall rule needs an IP range. That mapping lives somewhere across a dozen Terraform repos, and the bot has to find it.
- Format precision: Firewall rule syntax is unforgiving. A wrong protocol, a missing port variable, or a malformed option block causes a pipeline failure hours later.
- State management: A ticket doesn't have one state — it evolves. The bot opened a draft MR. The validate pipeline failed. A human edited the branch. The plan pipeline ran on a dev-to-main MR. Each of these requires a different action.
- Volume: At hundreds of tickets per year, even a ten-minute manual task per ticket adds up fast. And unlike a one-time automation project, access requests never stop arriving.
The architecture: a coding agent in a GitOps loop
The core idea is straightforward: take an AI coding agent — the kind that can read files, run shell commands, write code, and call external APIs — and give it a well-defined playbook for a single, scoped task. Not a general-purpose assistant. A specialist.
The bot is a Python orchestrator running a local open-weight language model (served via Ollama). It runs as a one-shot container: start, process all open tickets, exit. No daemon, no cron loop inside the process. The scheduler lives outside the container, which keeps the bot itself simple and easy to reason about.
The entire system runs on the engineer's existing workstation — no dedicated AI servers, no GPU cluster, no cloud inference bill. The language model runs locally via Ollama on the same machine that already has access to the internal network and source repositories. The company does not need to invest in any AI infrastructure to adopt this approach; the compute that matters is already there.
For each open access request ticket, the bot follows a state machine:
- No MR exists: Generate the firewall rule, push a branch, open a draft merge request.
- MR is open: Check the validate pipeline. If it failed, read the job trace, analyze the error, push a fix commit to the same branch. The MR updates automatically.
- MR was merged: Monitor the downstream pipelines — plan, then apply. If either fails, recreate the fix branch from the target branch HEAD and open a new draft MR. If apply succeeds, record a learning.
- MR was closed by a human: Skip. Human judgment takes precedence, always.
The bot never promotes a draft MR to ready. It never merges. It never comments on tickets. Humans make the final calls — the bot just handles everything up to that point, and cleans up after every pipeline failure so the human reviewer never sees a broken state.
Variable resolution: the hard part
Generating a firewall rule is easy once you know the destination IP range. Getting to that IP range is the actual engineering problem.
The bot follows a three-step resolution strategy for every destination named in a ticket:
- Check the existing variable definitions file. If the destination already has a named variable, use it. Add an evidence comment to the rule.
- Search sibling infrastructure repos. Grep for the service name across all related Terraform and infrastructure-as-code repositories. Parse the results to extract IP ranges, then create a new named variable.
- If undeterminable, use a placeholder. Add a
__PLACEHOLDER__variable, set a flag on the MR, and include a comment explaining what needs to be filled in. The human reviewer knows exactly what to do.
Critically, multiple unknown destinations are never collapsed into a single placeholder. Each gets its own, so that filling them in later requires no guesswork about which placeholder maps to which service.
The self-improvement loop
This is the part that makes the system interesting beyond just "automation."
Every time the bot fixes a pipeline failure, it captures what happened: the original rule it generated, what the pipeline rejected, what the human-corrected version looked like (if applicable), and the job trace. It calls the local language model and asks for a 1–3 sentence lesson: what was wrong, what the correct pattern is, and how to avoid repeating it.
That lesson gets appended to a learnings.md file. Then the bot calls the model again with a different prompt: given this lesson, which part of the step-by-step instruction file should be updated, and how? The model rewrites the relevant section and the bot patches the file in place, wrapping the changed block with attribution markers so the origin of every change is traceable.
The instruction file is re-read at the start of every subsequent run. So a lesson learned from a ticket on Monday is encoded into the bot's behavior by Tuesday, without any human touching a config file.
Two-tier review: fast learning vs. reliable learning
Local models are fast, cheap, and private — but they make mistakes. A lesson extracted by a 4B-parameter model running on a laptop can be subtly wrong, overly specific to one ticket, or in conflict with an existing rule. Running those lessons directly into production behavior is risky.
So we run a two-tier system:
- Tier 1 (fast, per-ticket): The local model extracts and encodes lessons immediately after each failure. Changes are marked with
[ollama-learned]attribution markers. These are live and active, but explicitly flagged as unreviewed. - Tier 2 (periodic, holistic): A stronger model — Claude Code, running on demand — reads all accumulated
[ollama-learned]blocks alongside the full learning history. It evaluates each one: is this rule correct for the general case? Does it conflict with anything? Does it address the root cause or just the symptom? It rewrites good ones, discards bad ones, and upgrades markers to[claude-reviewed].
Claude-reviewed blocks are trusted. Ollama-learned blocks are provisional. The distinction is visible in the file and guides future review sessions. The strong model never runs automatically — it's invoked intentionally, when the accumulated learnings warrant a pass.
Bootstrapping from history
When we first built this system, we had 300 historical tickets already resolved and merged on the main branch. Rather than starting the learning system from zero, we ran a one-shot bootstrap:
- Parse all existing configuration files for ticket references.
- For each ticket, retrieve the corresponding git commits on main.
- Extract the bot's original diff (what it generated) vs. the final merged diff (what the human approved).
- Call the local model to extract a lesson from each delta.
- Append all 300 lessons to
learnings.md. - Call the strong model for a holistic codebase review: given all 300 lessons, what should change in the instruction file and the code?
This transformed three years of implicit institutional knowledge into explicit, machine-readable behavior rules — in a single afternoon of compute.
The next step: task-specialized fine-tuning
The general-purpose local model is good enough to get started, but it is not optimal. It has no prior knowledge of this specific rule format, these specific variable naming conventions, or the patterns that appear repeatedly across hundreds of similar tickets.
The 300 historical examples are also a fine-tuning dataset. Each example is a (ticket description + access group + existing variable context) → (rule output) pair. The output format is highly consistent — typically 2–6 lines — making this an ideal target for LoRA fine-tuning on a small model.
We are using MLX-LM for training on Apple Silicon: LoRA adapters on a 1–4B parameter base model, fused after training, exported to GGUF, and imported back into Ollama. The fine-tuned model replaces the general-purpose one for rule generation only — lesson building and instruction patching continue to use the general model, where broader reasoning matters more than format precision.
The expected outcome: faster inference, lower hallucination rate on rule format, and correct variable naming without needing to see examples in the prompt every time.
Key design decisions
A few choices shaped the system significantly:
- Local model, no cloud calls for sensitive operations. Infrastructure data — IP ranges, service names, network topology — stays on the machine. The local model never sends anything to an external API.
- One-shot execution. The bot runs and exits. There is no persistent process to monitor, no in-memory state that can drift, no reconnection logic. Every run is a clean slate read from files.
- The fix branch always has the same name. If the bot needs to recreate a branch after a plan or apply failure, it uses the ticket key as the branch name, every time. This makes it trivially easy to find the bot's work and reason about its state.
- Attribution markers on all machine-written changes. Every line added to the instruction file by the learning system is wrapped in markers that identify who wrote it and when. Human reviewers can audit exactly what the bot has taught itself.
- Draft MRs only. The bot never decides that something is ready to merge. That judgment belongs to engineers who understand the risk context, the business priority, and the operational timing.
What this looks like in practice
After running through 300 historical tickets, the learning file contains lessons about protocol inference (when to use TCP vs. UDP vs. a generic IP rule), port variable naming conventions, how to handle services that expose multiple ports, when a hostname-based TLS match is more appropriate than an IP-based rule, and a dozen other patterns that previously lived only in the head of whoever handled these tickets last quarter.
New tickets — the ones arriving today — are handled by a bot that has seen 300 examples of what "correct" looks like, has extracted lessons from every failure in that history, and has had those lessons reviewed and refined by a stronger model. It is, in a meaningful sense, an experienced practitioner for this specific task.
The pipeline failure rate on bot-generated MRs dropped significantly after the bootstrap. Not to zero — new services appear, variable naming conventions evolve, the infrastructure changes — but the bot recovers from those failures on its own and learns from them immediately.
Conclusion
AI coding agents are most powerful when applied to tasks with three properties: the input is constrained and machine-readable, the output format is precise and verifiable, and there is an unambiguous feedback signal. Network access automation has all three. The ticket tells you what to do. The rule format tells you how to do it. The pipeline tells you if you got it right.
What makes this system more than just automation is the learning loop. A bot that runs the same prompt against every ticket will plateau at whatever quality level it starts at. A bot that extracts a lesson from every failure, encodes it into its own instructions, and periodically has those instructions reviewed by a stronger model — that bot gets better every week, without anyone writing a config file or redeploying anything.
The goal is not to replace engineering judgment. It is to handle the part of the work that does not require judgment, so that the judgment-requiring parts get more attention. Draft MRs, not merged PRs. Lessons, not decisions. A specialist that knows its lane and stays in it.