DarcFactory: The Dark Factory Era

AI writes code faster. Software delivery does not necessarily follow. When one agent finishes a task, a person checks it. When ten agents work at once, ten results arrive at that person's desk. Increasing production makes the other constraints more visible: organizing the work, verifying the results, and making the decisions that allow software to be delivered responsibly.
I believe the next bottleneck in software engineering is shifting from writing code to organizing, verifying, and supervising work. At ArcBlock, we call the production model we want to build toward DarcFactory: AI workers operating continuously within clear intent and boundaries, results receiving independent checks, and people able to see what is happening and intervene where judgment is needed.
A lights-out factory is a manufacturing operation that can keep producing without people continuously tending the floor. The machines do not need lighting for human workers, so production can continue with the lights off. The term describes unattended production; maintenance, exceptions, and human responsibility still matter.
Science fiction imagined an early version in Philip K. Dick’s Autofac, first published in the November 1955 issue of Galaxy Science Fiction. Its automated factories keep finding raw materials, making and delivering goods after a war, and can repair or rebuild themselves. The survivors want control of production back, but stopping the factories proves harder than keeping them running. The story leaves a question: once production becomes autonomous, can people still decide what it should do and when it should stop?
The name carries two meanings. Darc sounds like dark, recalling the manufacturing idea of a lights-out factory. DARC also stands for Decentralized Autonomous Resource Coordination, a protocol direction for independently owned Arcs, the computing environments used by people and agents, to coordinate work, resources, and verification. Our ambition goes beyond making one factory autonomous: independently owned computing resources could take on work and cooperate through verifiable results, gradually forming a collaborative economy.
During ASE 2026, I want to bring this vision into public discussion. We are exploring its mechanisms in our internal engineering practice. This essay sets out a direction for software production, with no announcement of product availability or open access.
ASE stands for Automated Software Engineering. It is a fitting place for this conversation. At its Harness4GenUI workshop, my keynote, live demonstration, and panel participation focus on AFS-UI: when interfaces are generated and models keep changing, what should remain stable? DarcFactory takes that question onto the production floor. Implementations can change continuously. The intent, boundaries, and evidence behind the work need a stable home.
The lights can go out. Judgment still has to happen.
The manufacturing metaphor is compelling: production continues without people constantly standing on the line. Applied to software, however, the absence of people is easier to picture than the conditions that make unattended work possible.
A production line needs a definition of acceptable output, a way to detect trouble, and a reason to stop. Software adds a difficulty: we often discover what we need while building it. A change that passes today's tests can make tomorrow's changes harder. AI workers produce nondeterministic output, and acceptance criteria can themselves be incomplete. Turning off the lights does not remove those problems.
There are serious experiments underway. StrongDM's software factory account, published in February 2026, centers specifications and scenario validation in an experiment without human code writing or review. OpenAI's Symphony connects a task board to continuously working coding agents while retaining human review of results. They explore different ways of moving beyond individual coding assistance into organized, continuing work.
The criticism deserves equal attention. In his AI Engineer talk, Harness Engineering is not Enough: Why Software Factories Fail, HumanLayer's Dex Horthy challenges a central assumption: passing tests on short tasks is an inadequate measure of long-term maintainability. A factory can keep finishing tasks while making the codebase harder to change. Adding more loops does not answer that concern.
I take that warning seriously. A factory should be evaluated beyond the number of changes it merges today. We need to ask what defects escape delivery, whether the next modification becomes harder, and whether people can understand and take over when the system encounters unfamiliar trouble. Independent review does not automatically solve these problems either. It introduces another judgment into the process.
That leads to a condition for our vision: autonomous production must rest on trustworthy verification. The reach of verification determines where autonomous execution has a basis to proceed. Work beyond that reach requires human judgment, or better verification before the autonomy expands.
Work needs an identity of its own
Much AI work still lives in chat histories. One terminal holds the requirement, one conversation holds the implementation, and another contains a review. A person carries context between them and remembers which opinion applies to which version of the code.
DarcFactory's design starts with the Work Object. A requirement, a defect, or a set of changes should have an addressable identity, with links to its intent, execution records, results, and verification evidence. Replace the agent or change the interface, and it remains the same piece of work.
Consider a bug fix. “Done” tells us very little. What was originally authorized? Which files changed? Which version was checked? Was the code modified after review? What conditions remain unmet? Those relationships need to travel with the work, rather than remain in one worker's memory.
DarcFactory's design relationship: shared records carry the work; people retain intent and consequential choices while execution and verification operate around the same object.
The work object provides a foundation for organization and a point from which people can observe it. People see progress in a cockpit; agents read and write the corresponding objects and records through AFS, the Agentic File System, an abstraction for file-system-style addressing and operations. The goal is a common account of the work. An agent saying it has finished is a statement. Whether checks passed on the relevant version and whether the result was delivered are facts that require their own verification.
This distinction matters. If “ready to deliver” is simply a field filled in by a worker, even a beautiful cockpit is displaying a tidier set of self-reports. We want status derived from evidence. When evidence is insufficient, the state should remain unknown or require attention.
An internal example I shared in August puts the work backlog, its age, and incoming and outgoing work on one page. It makes a practical problem visible: discovering issues can outpace resolving them. Seeing busy agents is not enough. People need to see where work is accumulating.

Internal example shared in an X post on August 31, 2026. The figures are a snapshot of the queue at that time.
Give workers autonomy. Check their results.
One principle in our factory design is responsibility for output, autonomy in process. The factory's central responsibilities are organizing assignments and arranging checks on results. Workers need room to solve problems and organize execution, within the task's authorization, permissions, and safety constraints.
Activity alone cannot tell the factory how well a worker is doing. Time spent and process counts are weak substitutes for delivered work, the checks it received, the rework it needed, and the problems found afterward. Those observations are closer to the capability we want to understand.
Here, an “engine” means an AI coding system such as Claude Code or Codex, rather than a new conversation within the same system. Our factory design uses cross-engine code review: the reviewing engine must be outside the set of engines that contributed to writing the result. One engine can write and another can check. The purpose is straightforward: introduce a different reasoning process before an implementation's misunderstanding gets confirmed by its own reviewer.
“Independent” has limits here. Different engines do not guarantee statistically independent errors. Models can share training material, assumptions, and blind spots. Cross-engine review is a structural way to increase diversity of judgment; its usefulness still needs to be measured through missed defects, rework, and actual outcomes. It belongs alongside deterministic checks, realistic usage scenarios, and human expertise.
Evidence must also apply to the version being delivered. Approval of yesterday's code cannot simply carry over after another change today. Passing checks cannot authorize changes outside the agreed scope. We want these conditions to become explicit, verifiable rules that support automated delivery within established authorization and acceptance criteria. New tradeoffs or new authorization return to people.
The early test-report example below distinguishes failures from incomplete checks and preserves the source and coverage of each run. Passing the checked portion cannot turn the unchecked portion into a pass. People need that distinction when deciding whether delivery can proceed.

Internal example shared in an X post on August 20, 2026. This is the test-record view from that stage.
Internal practice has reminded us that the factory itself makes mistakes. In one case, the system encountered an “out of quota” message quoted inside test code and misread it as an actual limit on an AI engine. Work stopped for several hours. Supervisory signals exposed the problem; it became a piece of work that went through repair and review by another engine. In another repair, the review engine found that an apparently correct change would also hide a class of real failures. The fix was sent back for revision.
Neither example proves that the factory can be left alone. They explain why production, checking, and supervision need separation. Producing a fix does not establish that the fix should be accepted. A system explaining why it stopped does not establish that the explanation is correct.
Supervision therefore needs signals that workers cannot unilaterally rewrite. Delivery, completed checks, and external operational health need observation from their respective sources of fact. Missing signals deserve attention; silence cannot stand in for health. That capacity to observe is a prerequisite for reducing routine human intervention.
AFS-UI makes work visible. DarcFactory organizes it.
This is why I want to connect the two conversations at ASE. AFS-UI concerns the relationship between an interface and shared operational context. The interface is a projection over a working environment. People interact through the rendered surface; agents access corresponding structures through supported addresses and operations. Interfaces can evolve while object identities and operational semantics remain understandable.
DarcFactory applies that idea to production organization. Work objects, reviews, and evidence need a shared operational home. A cockpit may provide many views without creating another account of the work for each view. AFS, the Agentic File System, provides file-system-style addressing and operations for that organization. A “file” here can be a resource or view accessed through a path, as well as a document on disk.
When a person sees blocked work, they should be able to reach its reason and evidence. When an agent takes it over, it should be able to read the corresponding intent and boundaries. These are the two sides we want to connect on the same foundation: understanding and intervention, execution and checking.
Our ambition is to move software production from workshops where people watch terminals into supervised autonomous production systems. The same organization of work may eventually be useful beyond software. But delegation in each domain depends on that domain's verification capabilities. A general abstraction does not justify a general expansion of autonomy.
There are substantial questions here for the ASE research community. Does less human intervention mean greater autonomy, or missed opportunities for necessary judgment? How much diversity does cross-engine verification need to reduce shared blind spots? How can long-term maintainability enter continuing acceptance checks? How should supervisors themselves be checked?
Internally, we are trying to record human interventions in successive rounds of work so that these questions become observable. Intervention counts need to be considered alongside task difficulty, rework, escaped defects, and subsequent maintenance. A factory that never asks for help and delivers bad results has no autonomy worth celebrating.
Another quality-dashboard example brings together issues found and closed, the queue needing human attention, and the proportion of tests actually executed. It also identifies gaps in coverage. What matters is the relationship between these observations: how much work has closed, whether people can keep up, and what verification actually covered. No single percentage in the picture establishes the quality of the whole factory.

Internal example shared in an X post on September 1, 2026; the dashboard is dated August 28. Its figures and coverage belong to that record.
I want DarcFactory to change the level at which people participate. People own intent: what is worth doing, which boundaries must hold, and what decisions to make when evidence is insufficient or tradeoffs arise. Machines carry forward the work that has clear authorization and can be verified. Human attention can move from repeated confirmation toward the places where it contributes judgment.
That is the dark factory era DarcFactory aims toward. Its measure is whether work retains a clear origin, results receive trustworthy checks, and someone can see, understand, and take over a problem even when people no longer stand continuously on the production floor. Under those conditions, turning off the lights begins to mean something.
References
- Justin McCarthy, Software Factories and the Agentic Moment, StrongDM, 2026-02-06.
- Alex Kotliarskyi, Victor Zhu, Zach Brock, An open-source spec for Codex orchestration: Symphony, OpenAI, 2026-04-27.
- Dex Horthy, Why Software Factories Fail; AI Engineer talk, 2026-07.
Referenced here
Events
Terms
- AFS
AFS (Agentic File System) turns the files, services, and active work relevant to a task into an inspectable resource view. Instead of handing an agent an undifferentiated machine or a pile of APIs, it gives the task a world with names and boundaries.