We're at an interesting point, aren't we? A lot of us are working with agents, building software and solving remarkably similar business problems. Accounts. Tax. Sales. Marketing. Purchasing. The things that keep a company running.

And I keep coming back to a question. If an agent can increasingly build the software, what is the valuable thing that we need to own?

I think it is our understanding of the business. Its data, its responsibilities, its agreements and the rules by which it operates. The code matters. But it may no longer be the most durable part.

The rules could become more important than the code that implements them. That is my argument, not a claim that we can already hand an entire company to a bot and go on holiday.

How many times do we need to build an invoice?

Think about how many accounting systems there are. Each needs to know who sold something, who bought it, what it cost, when it happened and how it should be recorded. There is a lot of common ground.

That does not mean every company's data is identical. Different countries, sectors, contracts and reporting requirements create real differences. But we do repeatedly build software around a substantial common core.

My first reaction is: surely we don't need to keep doing all of that?

Then I remember why competition matters. If everybody depends on one supplier and leaving is painful, that supplier has considerable power over the price. Another company comes along, does it differently or more cheaply, and gives customers a choice.

So I don't want one compulsory accounting system. I want something more useful: common, portable data and clear business rules, with competition over how well the work gets done.

The opportunity for agentics is not that businesses suddenly become identical. It is that implementing the differences might become much less expensive.

DevDay has made this much more concrete

I started writing this while waiting for OpenAI's announcements. Now the official DevDay recap is out, and there is something rather relevant to this argument.

Dots: ongoing work, with boundaries

OpenAI has announced dots: persistent agents powered by Astra, with their own cloud computer. They are rolling out to Pro and Business Premium in eligible markets; Enterprise, including Edu and Healthcare, has an administrator-enabled beta.

The important part for me is the boundary around the work. Custom Rules can permit actions, require approval or block them. OpenAI says background proactive research is read-only; that is different from authorised tasks that can act. Consequential work still needs review.

Specialist dots go further towards organisational roles, with their own identities and credentials. But these are a preview through focused enterprise pilots, not a generally available workforce. The first personal dot is included at no extra cost in Pro and Business Premium; deeper work has allowances, and Codex and Work tasks use their usual limits.

That distinction matters. I can ask something to help me all day. Giving it responsibility for part of a company is a different proposition.

Who may it contact? What may it promise? Which records can it change? When does it stop and ask me? Those are exactly the questions I mean when I say that the rules are becoming more important.

To me, this is a concrete step towards the direction I am describing. It is not proof that an agent can already run your business safely, or that writing a few instructions replaces proper controls.

Shared work needs shared authority

The Team Tasks documentation is revealing too. Team Tasks are available on Business and Enterprise plans, subject to workspace permissions. Scheduled or event-triggered work runs with a team's service account and configured connections. It does not inherit the creator's personal memories or chat history. A connected account can expose information beyond a member's own access, so membership and connection permissions need reviewing together.

There is a useful lesson there. The operating knowledge cannot live only in somebody's head, or in a private conversation with an agent. It needs to be made explicit, owned by the team and attached to the right authority.

The Agents API has a new capability, not a new birthday

The Agents API was already announced in public beta on 10 September 2026. Today's addition is computer use. Its documentation describes agents operating a browser in an OpenAI-hosted environment to interact with websites and applications.

That widens the implementation options. It does not decide which customer records your agent should see or whether it may approve an invoice. Those remain your decisions.

Grok Bot

Grok Bot is described as a beta offering of persistent agents with a shared cloud computer, able to work across existing applications and coordinate tasks. A supplier's examples do not prove reliable end-to-end operation in your company. Shared access also needs careful control.

But don't let the beta label make you dismiss it. I already see Grok Bot as a very powerful chief-of-staff-style assistant: helping organise the day, research questions, prepare work and keep things moving. With the right access and guidance, I think it can help with most of the things I want support with day to day. It still needs my judgement and oversight, but that does not make it any less useful. It is already a powerful source of support.

OpenClaw

This is where a working local pilot becomes really interesting to me. With OpenClaw configured and the right permissions in place, I can execute code on my own machine and choose from different supported models and providers, rather than be tied to one particular model. That combination is incredibly powerful: my tools, my working environment and a choice of intelligence to help me use them.

I would still need to get that pilot working and prove what it can safely do. Running code locally does not necessarily mean the model runs locally or that no data leaves the machine.

OpenClaw's security documentation covers access controls, tool permissions, sandboxing, browser risks and prompt injection. Installing an agent does not, by itself, create a secure or properly authorised worker. These product descriptions were checked on 29 September 2026.

The direction interests me more than the name on the button. We are moving from asking for an answer towards asking for a piece of work to be carried through. That makes the organisation of the work much more important.

Clean code. Clean for whom?

We put considerable effort into making code readable and maintainable. Quite rightly. Someone else needs to understand it, investigate a fault and change it without breaking everything around it.

But suppose agents do more of that implementation. Does the centre of attention move?

I think it does. I can see more human attention moving towards the data model, permissions, contracts between systems, the decisions we are trying to express and the evidence that the result is correct.

That does not make unreadable code a good idea. Readability also helps agents, audits, security reviews and recovery when something goes wrong. A tested, understandable system is still worth having.

What changes is what I would regard as the primary business asset. Not simply the current implementation, but the specification of what the business requires, together with the examples and tests that prove it.

If I can replace the implementation without losing the meaning, that gives me options.

From a board decision to something the system can do

A board might decide that the business needs better control of spending. Management then has to turn that into responsibilities, policies, agreements and processes. Eventually, software does part of the work.

The difficult part is not just translating English into code. It is discovering what the English actually means.

“Be sensible about spending” sounds perfectly reasonable until you try to build it.

An illustrative purchasing process, not a ready-to-use financial policy
LayerMake it explicit
Business intentControl spending without delaying legitimate purchases.
Policy and ownerA named finance owner sets approval limits and decides how exceptions are handled.
Data and agreementsIdentify the supplier, order, invoice, currency, relevant terms and authoritative records.
ProcessCheck for duplicates, match the evidence and prepare a recommendation. A mismatch goes to a person.
PermissionsThe preparing agent can read records and draft a request. It cannot change bank details or release payment.
ImplementationSoftware performs the allowed steps. Separate controls enforce approval before any payment action.
Evidence and reviewRecord the rule version, inputs, decision and approver. Test missing evidence, duplicates and attempts to exceed authority.

Now we have something that can be discussed, challenged and tested. We can ask what happens if an invoice arrives twice, if a supplier changes its bank details, or if the evidence is missing.

Those are business questions before they are coding questions.

A policy document is not a permission system

There is already an engineering approach called policy as code. For example, Open Policy Agent separates deciding whether something is allowed from the software that enforces that decision. This is not an invention that arrived with today's agents.

What interests me is bringing that discipline further into ordinary business operations.

Writing “do not pay an invoice without approval” into an agent's instructions is useful. It is not the same as preventing the agent from making a payment. The account permissions and the payment service need to enforce the boundary too.

Nor can every policy become a neat yes-or-no rule. Contracts can conflict. Laws require interpretation. Customers have unusual circumstances. Sometimes the correct executable rule is: stop, preserve the evidence and ask the responsible person.

I would keep the human-readable policy, the executable controls and the test cases together, versioned and reviewed. If they disagree, that is a problem to resolve, not something to let a model quietly improvise around.

And I am not proposing that the agent rewrites the payment system every time it sees an invoice. Generate a change, test it, approve it and release it through a controlled process. Cheap code is not a reason to abandon change control.

Let the system help choose the intelligence

There is another job I would like the harness to help with. Which model should do which piece of work, and how much effort should it use?

I do not think most business users want to spend their day choosing between Luna, Terra, Sol and Astra, then deciding whether this particular task needs high, extra-high or ultra effort. The names and available settings change. The business question usually doesn't.

I want the system to say: this is routine extraction; this needs a stronger check; this is ambiguous enough to require a person.

DevDay gives us a fresh reason to revisit those choices. GPT-6.1 Sol is available through the API and in ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu. It is not yet in ordinary Chat.

GPT-6.1 Sol standard API prices, checked 29 September 2026, in US dollars
Token typePer million tokens
Input$2.00
Cached input$0.10
Output$10.00

OpenAI describes its capability as close to Astra on several evaluations, at one-fifth of Astra's standard input and output token prices. That is a supplier's evaluation and a token-price comparison, not a promise that your complete workflow will cost 80% less.

The Decisions API announcement is relevant too: Luna answers developer-defined questions using a finite set of possible answers, for classification, routing or choosing a next action. It is in limited preview, with wider release planned, not already broadly available.

I can see that helping decide which work goes where. But selecting an answer from a list does not make it correct, and recommending an action does not authorise it. Keep the permission checks outside the model's discretion.

That is a real research area. RouteLLM studied routing between stronger and weaker models and reported substantial savings on its evaluated tasks while maintaining response quality. It supports the possibility of better cost allocation, not a guarantee for every business workflow.

An agent's confident guess about the best model is not enough. Measure the outcome. Include the retries, the routing overhead and the time people spend correcting it. Cheaper tokens are not necessarily cheaper work.

I would optimise for the cost of an accepted result, at the quality and risk level the task needs.

The economics will make people pay attention

This is where I think people will grow their businesses: getting more done with agentic support when they cannot yet afford another employee, once salary, employer taxes and the other costs of employing someone are taken into account. Not replacing a person they already have, but making room for work they could not otherwise take on.

The bookkeeper does not simply disappear

I can see more routine bookkeeping becoming automated: collecting records, proposing categories, matching documents and flagging gaps.

That changes the work. It does not prove that bookkeepers have ceased to exist, or that an accountant checking once a year will be enough. Some errors and obligations need attention much sooner.

The interesting role becomes understanding exceptions, checking the evidence, maintaining controls and helping the owner understand what the numbers mean. Similar questions arise in legal work, procurement and administration.

There is a distinction here that matters: a role is not just a list of tasks. Automating tasks does not automatically transfer responsibility, judgement or relationships to the software.

We still need the functions of an architect, a programmer, a reviewer and somebody coordinating the work. Some may be assisted by agents. They do not all require separate full-time posts in a small company. But the responsibilities still need owners.

What happens to software competition?

Here is the part I find particularly interesting. If your data, rules and tests are portable, you might be less dependent on the supplier that happens to implement them today.

You could ask another system to do the same work, then test whether it actually does. That is a stronger position than discovering that your business logic exists only inside somebody else's product.

But there is an obvious counterargument. We could exchange software lock-in for agent-platform lock-in. If the models, memory, identity, integrations and operating history all belong to one provider, moving may still be painful.

So portability has to be deliberate. Keep exportable records, documented interfaces and tests that are yours. Do not assume that generating code makes the whole system interchangeable.

I also care about sovereign AI and resilience. If we build essential work around agentics, we need credible alternatives when a provider, network or jurisdiction becomes unavailable. A server in Britain is not independence if its intelligence still depends on an overseas API.

For me, the practical questions are which models we can operate, where data goes, what can be exported, and how the business carries on during an outage.

That is also how I read DevDay's Private Intelligence announcement. It includes Zero Data Retention with Private Safety Processing, while Private Inference is a forthcoming autumn preview. Privacy protections are welcome, but they are not the same as sovereign control or an independent fallback if the supplier is unavailable.

This could be a very good time to understand how a business works

I think there is a substantial opportunity here for business analysts, operations people, risk officers and people working in governance and compliance.

Not because we need more documents that nobody reads. Because we need people who can turn an intention into something precise enough to execute, test and challenge.

Who owns the decision? Which record is authoritative? What is allowed? What evidence is required? Where must the system stop?

Those people will need to work closely with engineers and the people who actually do the job. The unofficial workaround may tell you more about the real process than the policy manual.

My prediction is that competition will make this move quickly in some markets. Not everywhere, and not because every company can safely automate everything. But if a competitor delivers the same reliable outcome with less effort, that is difficult to ignore.

For a small business, I would start with one recurring process. Describe the rules, collect a few real examples, define success and make the agent demonstrate it in a read-only or draft-only setting. Expand its authority only when the evidence supports it.

We have spent a long time treating source code as the thing we must protect and maintain.

I think the next question is bigger: can you explain how your company works clearly enough that another intelligence can help you run it?

That is not just a coding job. And that is what makes it interesting.