Skip to main content
Knowledge mapWhat Endures When Everything Evolves? — manuscriptArcBlock

You are here. See how this question connects to other ideas.

Select a node to open its page · Expand to read within the map

Knowledge mapFollow a connection. Understand a question.
← AFS and interfaces

Keynote manuscript

What Endures When Everything Evolves? — manuscript

Read the keynote argument, from the X Window inspiration to testable foundations for generative UI.

Open the presentation · 中文提纲与完整稿

AFS-UI and the Search for Stable Foundations in Generative UI

Draft for Robert Mao’s review. The first-person passages below use the design motivations and historical experience supplied in this task. Proposed evaluation questions and the wording of the argument remain editorial suggestions for approval.

Explore the Learning material behind each part of the talk

This reading index is outside the spoken manuscript. Numbers refer to the slides and sections below; follow a topic for examples, boundaries and technical sources.

SlidesArgumentContinue learning
03–04Unix and X Window influencesWhat did we learn from X Window? · Stable foundations learning path
05–08Resources, addresses and operationsStart with AFS · Working with resources
09–11Interface structure and two projectionsAddressable interfaces · Resource and view projection
12–14State, displays and sessionsDisplays and sessions
15–17Scope, composition and inspectionSmall World · Composition · Inspecting the structure behind a view
18AFS, AFS-UI, AUP and Web DeviceFrom protocol to website · The 18 core primitives · Web Device components
19Related protocols, SDKs and toolsComparison table · Comparison learning path
20–25Evaluation, limits and discussionRevisit stable foundations · Paths and semantics

01 · The question

We have become very good at asking what AI can change about an interface. Can it generate a screen? Can it adapt a workflow? Can it assemble the right controls for this person, in this moment? Those are useful questions. Today I want to add another one: when all of those things can change, what should remain stable?

My starting point is the experience of building AFS and AFS-UI in Arc. I will use that implementation to make the question concrete. The argument is about software architecture: how people, agents, data and displays can continue working together as individual parts evolve.

The presentation itself runs through the same website and presentation facilities we are discussing. That gives us something real to examine. It also gives us a useful discipline: we should be able to explain precisely what this running example demonstrates, and what would require another experiment.

02 · What changes when the screen changes?

Imagine a task called “Prepare the keynote.” It has a title, a status and an action that marks it ready for review. Today it appears as a card. Tomorrow, a model produces a better layout and the card becomes a row. Next week, the same work appears in a guided sequence.

I want four agreements to survive that redesign: which task this is, what its operations do, who may perform them, and how each view connects to its state. Those are the contracts I will keep returning to.

The person may need to learn where a control moved. But should an agent need a new way to identify the task? Should an integration need a different way to read its status? Should “mark ready” acquire a different effect because its button changed?

Existing APIs and multiple-view applications already address parts of this problem. Our experiment is to bring resources, interface structure and operations into a common addressable environment, so a participant can discover their relationship through the same system. Whether that reduces coupling is a question we can test.

03 · A memory from X Window

One of the experiences behind this design goes back to around 1993 or 1994, when I encountered X Window. What impressed me was very concrete: a program could run in one place and put its interface on a display somewhere else. A report did not have to remain attached to the screen of the machine doing the work.

In X terminology, the server handles the display and input, and application clients can run remotely. The naming can sound backwards if you first learned computing through web applications. But the separation is the interesting part. The place where computation happens and the place where a person sees and controls it are distinct architectural roles.

I am borrowing that question for a world with agents. What becomes possible when a display is treated as a system device, and the operational context is not trapped inside one rendering of the interface?

04 · Keep the boundary, reconsider the machinery

The lesson I take from X Window is a separation of responsibilities. We do not need to reproduce its protocol, transport or security assumptions to ask a similar architectural question today.

A new display technology can arrive. A new interaction style can arrive. The mechanism for describing a view can change. Through those changes, it is still useful to know where application work lives, how a display participates, and which operations a participant may perform.

That is what I mean by something enduring. It is a contract that allows other things to change with less disruption. A contract can itself evolve, but its evolution should be explicit and understandable to the parts that depend on it.

The same thought applies to the Unix influence on AFS. The value of a common interface is the work it lets different tools do together. Its usefulness has to survive contact with the differences between the resources behind it.

05 · Everything is context

AFS stands for Agentic File System. It uses a filesystem-shaped interface to organize resources that an agent or application can work with. Those resources may include files, services and parts of a running system.

The phrase “everything is context” needs a careful reading. It does not mean that everything is text, or that every resource should be copied into a model prompt. It means that the working environment can be presented as something a task can discover, address and operate on.

A prompt is one consumer of that environment. A human-facing interface is another. A tool performing an operation is another. Each needs access to the appropriate part of the environment, in an appropriate representation.

This is also a boundary on the claim. Organizing context does not solve cognition. A model can still misunderstand a resource, choose the wrong action or make an unsupported inference. The architecture gives those activities a more explicit place to happen.

06 · An address gives work a target

Keep our task, “Prepare the keynote.” For the architecture example, give it the illustrative address /work/tasks/keynote. Reading it returns its title and status. Its supported actions can be discovered beside the resource. A consumer can refer to the task without describing where its card sits on a screen.

The reference includes an environment as well as a path. Think of “this project’s workspace, /work/tasks/keynote.” A path copied into another workspace is not automatically a reference to the same task. We must preserve or translate that context when a reference crosses the boundary.

Within the chosen workspace, a redesign should not require changing that reference. Moving the task or replacing it is a different kind of change: the application must preserve an alias, migrate consumers, or explicitly report that the old reference no longer resolves.

So a slash-separated name is only the beginning. The useful agreement says what object it identifies, where it resolves and how long consumers may depend on it. That is the sense in which I am using stable addressability.

07 · Providers preserve real differences

A provider connects a resource implementation to the AFS interface. A mount places it within a namespace. For our example, a provider could expose the task collection under /work/tasks. The data might live in a service; mounting it does not copy that service into a local directory.

Now consider replacing that service. A compatible replacement must still return the agreed title and status, advertise the supported operations, and preserve their effects. If the old provider denies a viewer’s attempt to mark the task ready, a replacement that accepts it has broken the contract, even if both return valid JSON.

The shared interface also has to make differences visible. Perhaps the replacement supports reading but not writing. Perhaps it cannot send change notifications. A consumer needs that information before it assumes that an editor or a live view will work.

This is the provider boundary carrying meaning: common ways to discover and invoke operations, coupled with explicit capabilities and behavior. It offers a place to compare implementations. It does not excuse us from comparing them.

08 · Operations are part of the contract

For the task, reading obtains the current status. An illustrative “mark ready” action changes the status from draft to ready. These are different operations, even if a generated interface places their controls next to each other.

Suppose two people open the draft. One marks it ready while the other edits an older copy. The contract must say whether that stale edit is rejected, merged or allowed to overwrite the newer value. Where a provider supports conditional writes, the caller can supply the version it read and receive a conflict if that version is no longer current. The UI then has a meaningful recovery choice: fetch the new state and ask the person to reconcile it.

A read-only viewer’s rejected write is another distinct outcome. Retrying it is not the same remedy as resolving a version conflict. The provider and the authorization boundary must preserve these distinctions.

AFS provides a common operational vocabulary. This example specifies additional application behavior we would require of a task provider; it is not a claim that every provider implements identical concurrency rules. Changing a card into a row should leave those rules alone.

09 · What the human sees, what the agent addresses

Now we can return to the interface. A person sees the task title, its status and a control. An agent should be able to reach the corresponding objects, state and supported actions through their addresses.

That is the feature I most want to emphasize about AFS-UI. The agent should not have to reconstruct known application structure solely from pixels when the system can expose that structure directly.

Of course, agents already have other options. Browsers expose a DOM. Accessibility trees describe many controls. Applications expose APIs. The question is how the interface participates in the same operational environment as the rest of the system, so those connections do not have to be reinvented independently for every surface.

This is semantic correspondence. It does not promise a path for every shadow or decorative pixel. Images and visual reasoning remain valuable for understanding appearance. The point is to preserve a direct route to meaning that the application already knows.

10 · A small example makes the distinction visible

There are two examples here, with different purposes. The Learning page has a small local model: switch a task from a card to a row, and its illustrative addresses remain visible. That explains the idea quickly.

We also tested a real AUP session in the local runtime. AUP means Agentic UI Protocol: the language of interface nodes and their connections. The session contains a text node named demo-status. Reading its AFS tree path returns its type and the content that appears on screen. Calling the session’s patch action changes that content, and the connected browser updates without reloading.

We then changed the parent layout from a column to a row and back. The status node kept its identity and its readable path. This is a small, repeatable change experiment over actual interface structure.

Its scope matters: demo-status is a UI node, not yet a task service with business rules. We have observed node addressability and update delivery to one display. Connecting the node to our hypothetical task resource is a further application step. Keeping these two examples distinct lets us explain the design and show one implemented mechanism without claiming a complete task application.

11 · Context and projection have different jobs

There are two boundaries on this diagram. A resource projection determines the working resources available to a participant. A view projection determines how available resources are presented. I will use those full names to keep their jobs distinct.

For our task, the resource projection might expose its title and status to a reviewer while withholding the author’s private notes. The view projection might show that accessible task as a card or a table row. A visual redesign changes the second boundary; a change in the reviewer’s scope changes the first.

AFS-UI’s central idea is to treat the interface as a view projection over operational context. The task and its operations need not be redefined every time a model proposes a new presentation. The interface structure can itself participate in the addressable environment, as the real status-node experiment showed.

This makes a debugging question precise. Did the object change? Did the participant’s resource scope change? Or did only the view change? Each explanation sends us to a different part of the system.

12 · Decide which state belongs where

Our task’s ready status belongs to shared work. The scroll offset of a reviewer’s browser belongs to that display. Selecting the task could be personal state, session state or a shared selection, depending on the application. A common resource environment does not make those ownership decisions for us.

If we put the ready status into a separate browser variable on every display, we have created copies that need reconciliation. If we put everyone’s scroll position into one shared resource, we may make the displays fight over where to look. Both errors are easier to describe once we name the state and its owner.

The same distinction applies to this presentation. The presenter’s current slide can be shared with followers; an audience member’s private notes should remain private. A reader who leaves follow mode may keep a local page position.

These are ordinary application decisions with visible consequences. The foundation is useful when it gives those decisions explicit objects, operations and boundaries, rather than leaving state ownership hidden inside whichever component happened to render first.

13 · Display, session and connection

A display is where a view appears and interaction is received. A session groups state associated with an interaction context. A connection carries messages. Their lifetimes and responsibilities can differ.

Our local experiment exposed why the distinction matters. A second browser could open the same private session and receive its current view. A later patch reached the first browser but not the second. Opening the same session identifier twice was therefore insufficient to establish the audience-following behavior we wanted.

The runtime also has a separate live-channel mechanism. That suggests a different integration route, but we have not completed the browser-level verification for the chosen deployment. This is a useful implementation finding: the desired participation policy has to match the actual delivery mechanism.

For a user, the questions are simpler. Will my work survive reconnecting? Does this screen follow the other screen? May I browse independently? The architecture needs concrete answers to those questions, not just a diagram containing several displays.

14 · One presentation, several participation policies

Imagine connecting this room to one presentation activity. A follower reads the presenter’s page position. An independent reader keeps a personal page position. A participant in a poll submits a response through a voting operation. These are three different policies, not three names for the same connection.

We can describe their contracts before implementing the enhanced demo. Only the presenter can change the shared page. Followers receive that change and a late joiner receives the current page. A voter can submit an allowed response without gaining permission to edit the slides.

We can also specify a failure case. After a disconnection, the follower should recover the current page, not replay every intermediate transition as if the talk were starting over. The application must choose and implement that recovery behavior.

This multi-display application is the next experiment, rather than a result established by our private-session test. The useful architectural step is that delivery, state ownership and authority can now be tested separately.

15 · Shared context, different worlds

Return to the task workspace. The author sees the task and private notes. The reviewer sees the task but not those notes. AFS resource projections can select resources, change where they appear and restrict the operations made available.

Suppose both participants see /work/tasks/keynote. To claim that they share a task, the application must know that both references resolve to the same underlying resource. Matching path strings alone is not enough. If one view renames the entry, a cross-view reference needs the corresponding mapping. That is why our stable reference included its workspace from the beginning.

A bounded resource view is sometimes called a Small World in AFS. It lets a participant work with an appropriate portion of the system. The reviewer need not see everything the author can see in order to collaborate on the same task.

Addressability still does not grant authority. The execution boundary must enforce allowed operations. Likewise, restricting the visible resources does not make malicious content inside them trustworthy. Scope, authorization and interpretation remain separate responsibilities.

16 · Composition goes beyond the screen

Once resources and operations can be addressed consistently, composition can reach beyond placing components next to each other. A view can consume a resource supplied elsewhere. Another view can use that same operational object. A provider can change internally while preserving the contract its consumers use.

That is a useful architectural ambition because it lets the system evolve in parts. A model, a renderer, a resource provider and an interaction pattern should not always require a coordinated replacement of everything around them.

But composition has obligations. Two consumers may compete to update the same state. A reference may become stale. An operation may change between versions. Access may be revoked. A common namespace gives us a place to express these situations; it does not settle them automatically.

The practical question is what agreement each boundary preserves. If we cannot describe that agreement, we should be cautious about calling the system composable. Successful assembly on one occasion is weaker evidence than continued cooperation after a meaningful change.

17 · Look behind the running view

For the real session example, an inspector should let us follow one concrete chain. First, read the demo-status node and compare its content with the browser. Next, patch that node through the session action. Finally, observe the new content in both the tree and the connected display.

Those observations concern the AUP structure. The browser DOM answers a different question: how did this renderer turn the node into HTML? A content file answers another: where did the authored material come from? An inspector is useful when it labels which layer we are viewing.

For a complete task application, we would extend that chain backwards. Which resource supplies the status? Which action changes it? Where does permission get checked? Which update causes the displayed node to change? That is how an inspection tool becomes an explanation of the application rather than an attractive JSON panel.

Our long-term interest is to make this kind of examination available throughout the site. The small session experiment supplies a verified starting point; integrating a polished inspector into the presentation remains separate work.

18 · What this presentation establishes

Here is the architecture in four names. AFS organizes the resources and their operations. AFS-UI is the interface architecture over that working context. AUP, the Agentic UI Protocol, describes interface nodes, bindings and events. Web Device supplies the website machinery: content, routes, languages, themes and rendering.

The hypothetical task travels through these roles. A provider exposes the task in AFS. An AUP view connects a status node to the appropriate data and an interaction to the appropriate operation. A rendering environment presents that view. Web Device is the implementation we use for this website; the protocol vocabulary and the website’s theme components are different layers.

Our evidence today comes from two places. This native presentation and the linked Learning pages establish the site’s content and rendering path. The isolated session experiment establishes real node inspection, patching and layout change with a stable node path. Neither is a comparative performance result.

The further task-provider and audience-following experiments have been specified but not completed. That leaves us with a concrete implementation to examine and clear boundaries for the research claims we can make.

19 · Compare the boundary being standardized

I will focus on two related efforts here. A2UI describes declarative interfaces through streamed messages, separating structure from state and supporting path-based data binding. MCP Apps connects interactive HTML interfaces to MCP hosts through UI resources and host communication.

These are real architectural contracts. It would be inaccurate to say that other approaches offer only pixels or lack state and paths. The question is which boundary each contract governs.

For our task, A2UI helps describe a status display and bind it to data in the interface’s model. MCP Apps gives an embedded task interface a defined way to communicate with its host and tools. AFS-UI’s emphasis is the relationship between interface structure and a broader resource environment in which both can be addressed and operated on.

That difference in emphasis is not proof of superiority or incompatibility. An adapter might connect these approaches; its behavior would have to be examined. The Learning comparison includes the other SDKs and tools, but the question for this talk is narrower: when our task view changes, which contract does each participant continue to rely on?

20 · Make the stability claim testable

We have completed one small change experiment: the actual session layout changed from column to row, while the status node’s path and value remained intact. The consumer reading that node did not need to find a new screen position. That observation concerns interface identity under a layout change, not independence from every renderer or provider.

Now specify a stronger experiment using our task. Begin with a task resource, a card view and an agent that reads the status and invokes “mark ready.” Replace the card with a table row while keeping the resource provider unchanged. The agent should use the same resource reference and operation, and the operation should still produce the same status transition.

Failure is concrete: the old reference stops resolving, the action acquires another effect, or a read-only reviewer can suddenly mutate the task. A table that looks correct would not compensate for those failures.

Then change a different part: replace the provider, keeping the promised contract. Repeat the same task and denial cases. Separating the changes tells us which boundary held and which one introduced coupling. This is the experiment needed to support the broader stability argument.

21 · Measure the cost of maintaining those relationships

One possible evaluation would compare implementations of the same application under controlled changes. We could record the consumer changes needed after replacing a renderer or provider, failures caused by stale references, and the effort required to preserve operation semantics.

We could also evaluate agent tasks with different access methods: visual interaction, structured interface access and direct operational access where the application supplies it. Such a comparison would need the same task definitions and carefully described capabilities. It would be misleading to give one system privileged information and attribute the whole difference to architecture.

For multiple displays, useful measures include update behavior, recovery after reconnecting and correct separation of shared and private state. For access boundaries, denial behavior matters alongside successful operations.

These are proposed evaluation directions, not results I am reporting today. The implementation makes the questions concrete enough to study. Establishing comparative benefits requires a study designed around them, with baselines, failure cases and enough detail for someone else to reproduce the work.

22 · Addressability creates obligations

There are costs to making the operational world explicit. Someone has to decide which objects deserve stable identities, which addresses are temporary and how consumers learn that a contract has changed. Providers have to implement compatible behavior rather than merely similar method names.

An inspector can improve understanding, but it can also expose information if its access boundary is wrong. A shared view can simplify coordination, but it can also make ownership mistakes visible to several users at once. More explicit structure does not remove the need for careful design.

There are also domains where a visual representation carries information that a simple object model cannot capture well. A drawing, a photograph or a spatial layout may still need visual reasoning. We should expose meaningful structure where it exists, while keeping visual tools available for the work that requires them.

The decision is therefore about which stable abstractions earn their maintenance cost. An abstraction is useful when it preserves relationships that matter to the application, and makes the remaining differences easier to reason about.

23 · A learning system can expose the same idea

The linked Learning material keeps the detailed definitions, primitive catalog and comparison sources available after the talk. You can start with the question that interests you and follow its connections.

For the last part of the argument, return to “Prepare the keynote.” We started with a card, then distinguished the task resource, the interface node, the display and the participant’s authority. Those distinctions gave us four things to preserve and specific ways to detect when one broke. The Learning pages let us unpack any one of them without turning this talk into a reference manual.

24 · What should endure?

The task can move from a card to a row while its identity, operation effects, access rules and relation to the view remain agreed. Those are the foundations I want us to examine when we discuss Generative UI.

We saw a narrow piece of this in the real session: a node retained its path while its layout changed, and an addressed update reached its connected display. The task-provider scenario extends that mechanism into an application contract. The proposed change experiments would tell us how much of that contract survives more demanding substitutions.

AFS-UI is one concrete attempt to put these relationships into an inspectable operational environment. Its broader benefits should be judged through applications and comparative evidence, including the cost of maintaining those contracts.

The question is old and newly urgent. X Window helped me see computation and display as separate roles. Unix showed the value of common names and operations. Generative UI now makes the view unusually fluid. That gives us a reason to revisit those ideas and ask which agreements people and agents should be able to rely on while the view keeps changing.

25 · Questions to take into the discussion

As we discuss Generative UI, I would like to leave three questions open. When a model changes an interface, what contract should other participants be able to rely on? How should we test that the contract survives a change in renderer, provider or display? And how do we preserve useful shared context while giving each participant an appropriate view and authority?

These questions cross the boundary between interface generation and systems design. They also connect implementation work to research on usability, testing and long-term evolution.

The slides and Learning material provide a concrete starting point for that discussion. The next useful step is to inspect a claim closely enough that we can agree on what would prove or disprove it. Thank you.