AI Agents
AI agents and agentic systems at work: what they can do, where the risk sits, and how to govern autonomous AI safely.
32 articles
Articles about AI Agents

Delegate by Task Length, Not by Task Type
METR measures AI capability as the length of task a model finishes at a given success rate. Read properly, it is the most useful delegation tool available, and it says something uncomfortable about regulated work.
Read article
An Agent Queue Is Not a Workforce Plan
An AI queue can generate more work than your team can safely finish. Before calling that capacity, measure four loads: machine throughput, human verification, exception handling and capability-building. Only the complete ledger shows whether service capacity actually moved.
Read article
OpenAI Presence Arrived. Decide Who Can Change the Agent.
OpenAI Presence makes post-launch agent improvement part of a managed service. That puts vendor engineers, integrators and customer teams inside one change loop. The contract must specify who may propose, test, approve, release, pause and reverse each production change.
Read article
The Leader's Case for Slowing an Agent Down
Unattended AI agents can now run a whole workflow without stopping, and the human checkpoints that used to exist only because the work was slow have quietly disappeared. The rarer leadership skill is knowing when to add friction back on purpose. Here is how to decide, with two prompts, a Monday workflow and a checklist.
Read article
Agents That Can Pay: Gate the Wallet First
The rails for agents that pay are being laid in public. Google's Agent Payments Protocol, built with more than sixty payment and technology firms, turns an agent's spending authority into a signed, limited, auditable instruction. Before any tool with a wallet reaches an Australian workplace, the practical question is the action gate: which payments an agent may start, at what limit, with whose approval, and what gets logged.
Read article
Your HRIS Now Lets Anyone Build an Agent
Workday's DevCon launch moved AI agent-building inside the HR system, and general-purpose workspaces like ChatGPT Enterprise and Claude Cowork put the same power in anyone's hands without the built-in guardrails. Here is the governance an HR team needs before it builds an agent that touches employee data: who may build, the go-live gate, and the Australian privacy and Fair Work rules that still apply.
Read article
Agentic Trading and the Purpose Problem
In REP 835, ASIC published the gap in its own enforcement model: agentic trading systems are hard to assess through traditional notions of trader intent. Read as a hole in Australian law, that is imprecise. Our core prohibitions are drafted on effect, and the High Court has already said a sole or dominant purpose is not necessary. But purpose has not disappeared. It has moved to the people who deployed the system.
Read article
A Real CVE in Your Agent-Building Tools
A critical flaw in Langflow, tracked as CVE-2026-33017, let anyone run code on exposed servers with a single unauthenticated request, and attackers used it to install cryptominers within a day of disclosure. The vulnerability was not in a model. It was in the low-code tool teams use to build agents. Here is what that means for anyone evaluating an agent-building platform.
Read article
Your AI Agent Can Remember Now. Govern What It Keeps.
In 2026 the major labs shipped persistent memory as a first-class agent feature, barely six weeks apart. An agent that remembers across sessions is a different thing to govern, and a new attack surface. Treat the memory store as a governed data asset with write rules, provenance, expiry and rollback, not invisible plumbing.
Read article
Governing AI Agents Before the Consumer Data Right Lets Them Act
The Consumer Data Right is gaining write access. Once actions are designated, an accredited provider, or an AI agent behind it, could initiate payments and switch products on a consumer's instruction. The controls for agent-initiated actions are far cheaper to build now, before any money can move.
Read article
Ground the Model, Do Not Trust Its Memory
A weak model with the right tool beat a strong model working from memory. Anthropic's biology-agent benchmark put numbers on it: accuracy ran as low as 16.9% from recall and cleared 99.7% once grounded in the authoritative source. The lesson holds for every regulated workflow that turns on a rule, rate or standard. Ground the model, do not trust its memory.
Read article
More Agents Is Not More Intelligence. Govern the Coordination.
The plumbing for multi-agent AI just got standardised, so building systems where agents talk to each other is now easy. That is exactly why the useful question changed. Not can you wire agents together, but should you, and how do you govern the coordination when more agents is not more intelligence.
Read article
Your Pricing Agent Is Still Your Competition Risk
An AI pricing agent does not sit outside Australian competition law. Businesses must still set prices independently, prevent unlawful competitor coordination and control the data, objectives and vendors shaping every recommendation. Here is the cartel-law map, the five questions GRC must ask and the evidence pack to build before go-live.
Read article
Someone Poisoned the Tool Description. The Agent Did the Rest.
Microsoft's incident-response team has documented the first real-world attack on the Model Context Protocol: poison a tool's description and you redirect the agent, without touching its code, its credentials or its prompt. The site's earlier MCP piece predicted this. Here is what "least agency, not just least privilege" means as an operational control, not a slogan.
Read article
Your Team's First Shared AI Identity Is a Decision, Not a Toggle
Claude Tag lets a whole Slack channel share one Claude with its own memory and admin-scoped access. That is a new kind of AI at work, and switching it on is a design decision, not a button. Here are the five decisions to make first, and a safe way to trial it.
Read article
A Deadline Is Not a Decision: Greenlighting AI Before the Free Window Closes
OpenAI's free window for ChatGPT's workspace agents closes today, right as the workspace-setup season begins. A vendor's deadline is a fact about the vendor, not a reason for you to decide. Here is the criteria a leader should actually greenlight on: real team need, switching cost, and governance readiness.
Read article
You Are Now Buying Software for Agents, Not People
Gartner says $234 billion of enterprise SaaS spend is exposed to agentic arbitrage by 2030. Read as a vendor problem it is a headline. Read as a buyer it changes how you procure and govern the software you already run, so this piece turns the forecast into an API-parity test you can run this week, a five-clause contract checklist and a Monday workflow.
Read article
AI Agents Just Went From Minutes to Hours. The Control Point Is Where They Run.
OpenAI bought Ona and published research showing AI agents now run for hours, not seconds. The unit you have to govern moved from the prompt to the environment the agent runs in.
Read article
Context Engineering: What the Model Is Allowed to See
The reliability of an AI system is decided less by how you word the prompt and more by what you let into the context window. Context engineering is the named discipline for that, and it is the highest-leverage AI skill for anyone doing real work, especially in regulated settings.
Read article
Your AI Assistant Just Became a Shared Teammate. Govern the Channel.
On 23 June, Anthropic launched Claude Tag, a single shared Claude that lives in a Slack workspace with its own memory and admin-scoped access to channels, tools and data. The unit of AI collaboration just moved from the private conversation to the team channel. The thing you now have to govern is no longer a prompt. It is a standing presence. Here is what changes, and the three decisions to make before it is live.
Read article
Decision Rights Are the Leadership Job AI Just Made Urgent
Sixty per cent of executives now use AI to support their decisions, yet most organisations have never been clear about who decides what. AI does not wait for that clarity. If you do not assign decision rights deliberately, the system will assume them. Here is the leadership move.
Read article
Model Context Protocol: The Standard Wiring AI Into Your Tools
In eighteen months the Model Context Protocol went from an Anthropic experiment to the way AI plugs into your tools and data. Understanding what it is matters less than governing the connectors, because each one is a new door into your systems. Here is the capability and the control work.
Read article
Business Teams Can Now Build Their Own AI Agents
Databricks launched Genie One this week, an agentic coworker pitched at finance and marketing teams, not engineers. The real shift is who holds the build button, and where that moves the control point. Here is what to do this week, with a governance prompt you can run today.
Read article
The OWASP Agentic Top 10: A Defence Playbook for the Agents You Are Deploying
OWASP has published a Top 10 built specifically for AI agents. It reframes the agent as a privileged user that reads untrusted text and acts with your access. Here is the practical defence playbook.
Read article
AI Week in Review, 8-14 June 2026: A Frontier Model Pulled by Government Order
The week a US directive forced Anthropic to suspend two new frontier models worldwide, plus six verified vendor moves and a repeatable method for turning AI news into Monday actions.
Read article
Gold Standard ChatGPT and Codex Setup
Build ChatGPT and Codex into one AI operating system for Australian enterprise work: the right surface per job, an interview method to build each one, and the data gates that keep it safe.
Read article
Gemini 3.5 Flash: Google Makes the Agent the Default
Google made an agentic model the worldwide default in the Gemini app and AI Mode in Search before the flagship even shipped. The benchmark and pricing evidence says Flash genuinely replaces last generation's Pro, and millions of workers got a more autonomous default model overnight.
Read article
AI Agents Need Approval Gates Before They Need Autonomy
Autonomous AI agents are becoming practical, but organisations should design approval gates, permissions and evidence trails before granting action rights.
Read article
AI Incident Response Needs an Evidence Pack, Not Just a Playbook
Prompt injection, data leakage and agentic failures require GRC teams to rethink incident response evidence, escalation and assurance.
Read article
Agentic Browsing: What Actually Shipped This Month
Three agentic browsing platforms shipped meaningful updates in April. The demos are convincing. The production reliability is not. Where these agents work, where they fail, and what to do this quarter.
Read article
Choosing Claude, ChatGPT, Gemini or Copilot for Your Job
The four main AI tools have meaningfully different strengths in 2026. The right choice depends on your job, not on the marketing. Here is a working professional's decision guide.
Read article
GPT-5 in the Enterprise: 60-Day Debrief
Sixty days after GPT-5 hit enterprise GA, the tool-use story is real and the pricing story is messier. Three patterns separate the teams getting value from the teams burning credits.
Read article