How I’m securing my agents · Part 1

Securing my agents starts with where they run

I already run always-on agents. Now I want an environment I control and a clear view of where each protection applies. My first test starts with a microVM inside VMware.

Editorial illustration of a small task workspace inside nested computer environments, connected to two separate compute systems.
Illustration generated with Gemini. The real experiment captures appear below.

I already have agents running around the clock. What I want now is a clearer picture of the environment they operate in and exactly where I am protecting it.

Where does each agent run? What can it access? Which control enforces that limit? Where can I see what happened when something goes wrong?

I want an environment I control, with protections I can point to and test. That is the build I am documenting in this series.

I already have VMware, the hardware, and GLM running locally on my two NVIDIA DGX Sparks. I want to use those pieces to understand the full path from an agent's task to the systems it touches.

That is why I do these lab projects. When the runtime fails, the browser stops, or the model cannot answer, I want evidence I can inspect and a system I can repair. Owning the environment makes those operating decisions my responsibility. It does not prove the controls work.

This is Part 1 of How I'm securing my agents. I am starting with where execution happens. The first question was narrow: could a small guest inside my VMware lab run a task against GLM on my DGX Sparks?

Yes. A Firecracker microVM booted inside an Ubuntu VMware VM, and a Python task in that guest received a valid response from local GLM. That establishes a working execution path. It does not mean I have moved all my agents into microVMs or secured this one. Permissions, network enforcement, and recovery are the next tests.

Give the agent a computer. Keep the model on the DGX.

An agent needs somewhere to work: files, a runtime, and tools. The model needs compute and memory for inference. I want to manage those parts separately.

Firecracker runs a small virtual machine, often called a microVM. In this experiment, that guest had 1 vCPU and 512 MiB of RAM. GLM stayed on the DGX. I did not need to put a copy of the model inside the microVM.

The compute stack was: Physical ESXi host -> Ubuntu VMware VM -> Linux KVM -> Firecracker microVM -> Python task.

Diagram of the nested VMware and Firecracker task environment and the separate model path through a temporary Mac bridge to local GLM.
Figure 1. The path tested in this experiment. This is an explanatory diagram, not a console capture or a validated security boundary.

KVM is the Linux interface that Firecracker uses for hardware virtualization. The VMware guest needs to expose that capability so another guest can run inside it. That is nested virtualization.

The model request went through a proxy for this task and an SSH management bridge to the existing model endpoint, then to GLM on the DGX Sparks. The guest did not receive the model credential. That credential stayed outside the guest.

The temporary SSH forwarding path passed through my Mac to reach the existing model endpoint. That let me prove the task worked, but it would tie an unattended agent to my management computer. A persistent private model path is one of the next things I need to build.

The useful design choice is separating the agent's working environment from the machine serving the model. I can change the guest runtime without moving the model into it. This test used my VMware and DGX lab; the separation is the part I want to carry forward.

Start with the VM you actually have

Hermes is the agent runtime I already use in a separate VM. That VM had no vmx or svm CPU flag and no /dev/kvm device. Firecracker needs KVM access. An agent already running in a VM does not mean that VM is ready to host another one.

I used a clean test VM instead. It had 2 vCPUs, 4 GiB of allocated RAM, and a 24 GiB thin-provisioned disk. Nested virtualization was enabled. The live Hermes VM stayed unchanged.

VMware CPU setting with nested hardware virtualization circled.
Figure 2. Actual VMware CPU setting, circled for inspection. Nested hardware virtualization is enabled on the clean test VM.

That distinction matters for this story. The inner guest ran one bounded Python task. Full Hermes, its browser, and its scheduler were not inside it. I tested the small agent computer first.

Three starts. One real model task.

I used Firecracker 1.17.0, a Linux 6.18.51+ guest kernel, and an Ubuntu 24.04 root filesystem from Firecracker's official CI artifacts.

I started a fresh Firecracker process and guest three times on the same Ubuntu host. The host's caches were not cleared between runs.

These timings measure when an SSH command succeeded. They do not measure kernel boot time. RSS is the resident memory of the host-side Firecracker process, sampled from /proc at that point in the outer Ubuntu VM.

Recorded console with the first guest reaching SSH readiness in 4.818 seconds circled.
Figure 3. Original trial 1 console: 4.818 seconds from process start to successful SSH. The annotation points to the first measured result.

The 512 MiB guest allocation, the process RSS, and the outer VM's 4 GiB allocation are different numbers. The outer VM still counts. Its memory was allocated, not reserved. I have not measured the whole stack's memory use or made a matched comparison with another deployment.

Trials 1 and 2 tested startup only. Inside the third guest, the Python runner read a controlled fixture and asked GLM to summarize it. The returned answer described the nested Firecracker setup and the separate DGX model server. It did not make up cost, uptime, or performance results.

The model identifier returned by my local API was GLM-5.3-Flash-EXL3. That is the local serving label. I did not inventory the model weights, quantization settings, or model-server memory use in this experiment.

The model request completed in 3.525 seconds, including forwarding and the response. It reported zero cached prompt tokens. This timing is separate from guest startup. These article measurements come from the original run; the later console recording has its own timing and cache record. I saved the answer, guest identity, and timing. This is one request measurement, not a throughput benchmark.

Recorded console showing the fixture summary returned by GLM and a 3.525-second round trip.
Figure 4. Original task in trial 3: one controlled summary request. The 3.525-second value includes forwarding and the model response, not guest startup.

The guest also had no default route. A route lookup for 1.1.1.1 returned "Network is unreachable." That confirms the missing default route. It does not test every possible network path or establish a complete network sandbox.

What I need before trusting this environment

The next version should stay up without the management-computer bridge. It should run repeated tasks, save its work, and recover after a model outage or restart. I want task records and measured recovery, not just a successful demo.

I also need to harden execution. This run used administrator access, no Firecracker jailer, and a CI guest image. Firecracker's guide treats its starter resources as demonstrations and describes a jailer for secure production execution.

There is a VMware support boundary too. Broadcom generally excludes nested hypervisors from support outside its listed limited Hyper-V cases. This experiment is a lab result, not a supported production design.

I have not completed a 24-hour reliability test or proved a cost or security advantage. Hardware, power, administration, and licenses still count. Those are things to measure as this becomes useful.

Next: own the controls as well as the computer

For the agents I already run, I want the controls to be explicit. I need to decide which identity it uses, which systems it can read, which actions it can take, and what happens when a request should be refused. A local model does not answer those questions by itself.

Controlling the execution environment is step one, and this experiment is the start of it: a small guest I can start, stop, inspect, and rebuild, with the model kept outside it. The next step is controlling what the agent does inside that environment.

That is where I want to take the series next, with AgentMinder. I will record the actual console and follow a request through the controls. I want to show an allowed action, a denied action, and the record that explains each result. Those are planned tests, not results from this microVM experiment.

After AgentMinder, I want to test the network boundary. The agent needs to reach the model, but that does not mean it should reach every workload in my lab. I will examine where NSX and vDefend can enforce that separation and test what they actually see in this nested setup.

Then I will look at what I bring into the environment: model files, packages, and tools. That includes how I select and verify downloads from places such as Hugging Face. A local file still needs a reason to be trusted.

Finally, I want to interrupt the model connection and restart the runtime. Does the agent preserve its work? Does it recover within its permissions? After those checks, I can run a measured unattended trial. That is how I want to test reliability in this environment, including whether its controls still hold during a failure.

Owning the pieces gives me somewhere to enforce those decisions. It does not remove dependencies: I still use VMware, Firecracker, model software, and the hardware underneath them. The goal is to reduce reliance on someone else's hosted agent service and keep the operating controls with me.

The first result is small and concrete: a 512 MiB microVM inside my VMware lab completed one task against local GLM. I saved the console evidence and the measurements. Now I can work through each protection around that environment and show where it applies, what it blocks, and what evidence I have that it works.

References

Field notes

Keep up with Field Signal.

New engineering field notes and coverage from the automated desks, in one email list. Unsubscribe any time.

Follow Pranav on X: @Pranav_patel06