Homelab / Private AI / Field notes

How I moved the Field Signal newsroom into VKS on the Dell.

The build behind the workflow: Avi, private networking, persistent storage, the cutover, and the tests that showed the workers could publish a story.

Illustration of a workstation, container workloads, routing, and a compact local AI computer.
Illustration: Field Signal, generated with Google Gemini. The screenshots show the actual lab.

Build record: This article documents the October 7-8 migration and its Qwen/Clef acceptance tests. The live newsroom now selects GLM-5.3-Flash-EXL3, and Clef is retired from the active workflow. The historical tests below describe the configuration used at that time.

The Field Signal rebuild gave each machine a job: MS-A2 coordinates the work, Dell runs it, and Spark serves the models. This is the installation record behind that layout. It covers the parts that had to work before the newsroom could move.

The goal was to run the application as VKS workloads on the Dell while keeping model inference on the Spark. The migration also had to preserve drafts and task history, keep model access private, and avoid two publication schedules running at once.

MS-A2 coordinates tasks, Dell executes the newsroom workloads, and Spark supplies local inference.
Workload placement. The Spark now serves the configured GLM model to the newsroom; Qwen and Clef were used in earlier migration checks.

MS-A2 is the brain. Dell is the workhorse.

The hardware had different jobs available to it. I wanted to use that capacity instead of putting every new service on the management machine.

MachineWhat runs thereWhat it does for Field Signal
MS-A2Paperclip, its management environment, and newsroom reconciliationKeeps the task view and coordination together.
DellVKS newsroom workloads, the private newsroom API, and the Spark ledger collectorRuns the workers that move stories through the pipeline.
SparkLocal inference; Qwen and Clef during migration, GLM now selectedSupplies local model responses to the newsroom.

Paperclip is where I can see the work and review personal articles. The Dell executes the newsroom jobs. The Spark answers the model requests. A task appearing in Paperclip does not mean the model is running there.

The Dell's guest Kubernetes cluster has two node VMs. The control plane has 2 vCPUs and 8 GiB of memory. The worker has 8 vCPUs and 24 GiB. Both reached Ready. The control plane manages the cluster, and the worker runs the application containers. Routing and Avi also have supporting VMs on the Dell.

The newsroom worker capacity in vCenter.
The worker VM has 8 virtual CPUs and 24 GB as displayed by vCenter. The utilization shown is one moment in time.

There is still only one physical Dell underneath this. If that host goes down, newsroom execution goes with it. This layout separates responsibilities, but I have not tested it as a high-availability system.

Getting VKS ready took some work

I wanted the agents to run as proper workloads in the lab. VKS, VMware vSphere Kubernetes Service, gave me a way to do that inside the vSphere environment I already use.

The installation was not just a container deployment. The selected networking path required Avi. The Dell also needed the right network connections, an accessible VM image, and enough resources for the supporting services.

The final setup works with a single physical uplink. A dedicated router VM connects the private management and workload networks to the lab services they are allowed to reach. That router runs on the Dell. The private port groups do not each need their own physical Ethernet cable.

Avi runs a controller and two service-engine VMs. The controller manages the configuration. The service engines handle load-balanced traffic. Getting them deployed meant obtaining the controller image and correcting the permissions Avi needed in its placement scope. Both engines eventually connected, and Kubernetes API traffic and private application traffic passed through the setup.

The two Avi service-engine VMs in the vCenter inventory.
The two Avi service-engine VMs in vCenter. Their connections and traffic were checked separately from this inventory view.

There were a few useful lessons in the failures. The original VM image lived on a vSAN datastore that the Dell could not access. Copying it into a Dell-local library let the image cache become Ready. The Dell Supervisor needed a supported resize, after which its running guest had 4 CPUs and 16 GiB. An earlier OutOfcpu event pointed to CPU scheduling, not exhausted memory.

Guest bootstrap also timed out before the Avi path was ready. Recovery used the supported machine remediation path, and replacement nodes reached Ready. API timeouts and high I/O wait showed up during provisioning too. That was evidence to investigate storage latency, not enough to name a failed physical disk.

This is part of why I wanted to build it. A model endpoint can be healthy while the application still cannot reach it. A VM image can exist while the host that needs it cannot access it. Working through the full setup means learning the dependencies behind the dashboard.

Moving the drafts without losing the history

The newsroom uses a 20 GiB persistent volume on Dell VMFS storage. It is configured for the current single-worker placement and has a Retain reclaim policy. Retain helps prevent automatic reclamation in the relevant deletion path. It is not a backup.

The Dell VMFS storage policy reports compliance.
The worker VM's storage-policy view. The newsroom volume was also checked through Kubernetes.

The original schedules stayed active while the replacement was being prepared. During cutover, the source was stopped and a final state copy was taken. Verification checked 4,474 restored files with zero mismatches. The old editorial timers and separate sports cron entry were retired before the replacement publication schedules were enabled.

A consistent SQLite backup was also copied off cluster and passed an integrity check. Scheduled off-cluster backups are still something to finish. Keeping the drafts and task history was part of the move, not cleanup to deal with afterward.

The model helps. The checks still matter.

The Spark connection is private, authenticated, and certificate-verified. Requests with missing or incorrect authorization returned HTTP 401 during acceptance checks. The target containers use pinned images and run without root privileges. Only Publisher receives the repository deployment key.

Reviewer uses Qwen and Clef, but they have different constraints. Clef has a smaller context allowance, so its adapter checks chunks of an article. Early context-limit failures were kept in the record before the adapter was corrected. Qwen separately reviews the full captured source material.

Local models still make mistakes. An earlier five-request experiment completed, but only three of four objective checks passed. JSON extraction and an editorial word-count requirement failed. A response coming back is only one part of the test.

The same applies to this article. Model review helps find problems. It does not turn a claim into a fact. Screenshots, source records, and actual tests have to support what I say happened.

Getting all the way to a live story

The pipeline had to do more than produce a draft. Six workflow tests covered exclusive claims, stale leases, bounded retries, publication holds, private source-address rejection, and exact evidence quotations. Real Qwen and Clef calls also passed from pods on the Dell.

The first new Publisher run still failed. Site validation picked up a temporary acceptance checkout. Moving that checkout under the existing output exclusion allowed the production build to pass. Then Vercel blocked deployment because the automation commit email did not match a GitHub account. Restoring the repository's existing deployment identity fixed that without rewriting history.

The successful build checked 410 public pages and 573 responsive candidates. Three stories then reached the live site: a markets story, a Volkswagen ID.4 recall story, and a Ligue 1 rights story. All three returned HTTP 200 with article content, and the story store marked them published. All four fresh Paperclip role tests succeeded too.

Sports, cars, and markets each have a daily discovery pass for at most one new story per desk per UTC day. The sports story went through a revision before release. My personal articles still need my approval. Newsletter and social delivery remain off.

That is the division I want. The routine desks can work through a defined process. My personal writing comes back to me, because the point of Field Signal is still my experience and what I learn from it.

Paperclip had to get out of the way too

A slow task view makes the whole system harder to use. Two problems needed fixing after the newsroom was running.

The synchronizer was comparing a shortened task-list preview with the full draft. It treated the difference as a change and rewrote all 24 stories every minute. A revision fingerprint fixed that comparison. The next unchanged reconciliation made zero writes.

A comment request was also searching 4,203 historical agent runs during secret redaction, with one branch missing a database index. Adding the index cut that measured database query from about 4.15 seconds to 0.4 milliseconds. The same authenticated comment requests improved from 4.5-7.7 seconds to 0.10-0.28 seconds. Secret redaction stayed in place.

Those measurements are for specific requests. They do not describe every screen or the speed of the models. But they address a practical problem: I need the place where I review work to respond when I use it.

How the workers run

The four newsroom roles use CronJobs that check for work every minute. Each run handles at most one story. The newsroom API and ledger collector use Deployments to keep application instances running. The containers execute in pods on the Dell worker node.

Durable story records, transactions, deduplication, and expiring leases coordinate work. Execution errors retry up to three times before becoming blocked. Retries reuse the article ID and URL slug. These controls reduce duplicate work, but they do not prove exactly-once behavior for every external publication action.

What this build establishes

The replacement nodes reached Ready. Restored state matched the checked source files. Local model calls passed from Dell pods, and three stories completed publication. Those are the acceptance results for this migration.

There is still one physical execution host. Automatic off-cluster backups are not configured. Avi evaluation licensing needs a continuing plan. Cluster upgrades, host failover, and sustained throughput have not been tested.

VKS gives the newsroom a consistent workload definition, job history, and persistent storage inside the vSphere environment. The installation still depends on working networking, image placement, permissions, and storage. That is the useful distinction from this build: the platform simplifies running the application once those prerequisites are in place.

References

Field notes

Keep up with Field Signal.

New engineering field notes and coverage from the automated desks, in one email list. Unsubscribe any time.

Follow Pranav on X: @Pranav_patel06