San Francisco Daily 360

collapse
Home / Daily News Analysis / Selector brings Git-based workflows and incident replay to network AI agents

Selector brings Git-based workflows and incident replay to network AI agents

Oct 11, 2026  Twila Rosenbaum 15 views
Selector brings Git-based workflows and incident replay to network AI agents

Key facts

  • Selector Foundry is a development and runtime environment for network operations AI agents.
  • Agents run on Pydantic AI and are defined as configuration, similar to infrastructure as code.
  • Customers commit agents to their own Git repository and use existing pull-request workflows.
  • Before production, agents are replayed against historical incident data and compared with recorded outcomes.
  • Failed promotions roll back in a single step.
  • Guardrails cap cost and token usage and constrain what an agent can conclude.
  • Full autonomy is the goal, but adoption is expected to be gradual and asymptotic.
  • Common workflows include cloud outage troubleshooting, provider maintenance checks, ticket filing, and status updates.
  • Future priorities include agent interoperability, MCP and A2A protocols, and monitoring for neoclouds.
  • Selector differentiates by letting customers build their own agents rather than choosing from a fixed vendor set.

Network operations teams are being asked to hand real decisions to AI agents. Building an agent that can analyze data is one challenge. Trusting it enough to let it act in production is another. The gap between a promising prototype and a dependable operational system is where many AI projects stall. Software engineering faced a similar trust problem years ago and solved it with version control, peer review, automated testing, and staged rollouts. Selector is applying that same discipline to AI agents in network operations with a new technology called Foundry.

Selector Foundry is a development and runtime environment that lets network operations teams build, test, version, and govern their own AI agents inside the company's platform. The vendor develops a NetOps platform that was updated earlier this year to provide a correlated view across branches, colocation facilities, on-premises data centers, and public cloud infrastructure. The goal is to move from telling operators what is wrong and where to find data to helping them decide what to fix and how to fix it.

That shift matters because network operations is inherently cross-domain. A single incident can involve a cloud provider, a transit link, a branch router, a data center switch, and an application dependency. Operators often have to stitch together evidence from multiple tools, then follow manual playbooks that vary by team, vendor, and severity level. AI agents promise to reduce that toil, but only if they can be trusted to follow policy, avoid hallucinations, and produce repeatable outcomes.

How the platform works

Foundry treats an agent the same way a team would treat a piece of software, from framework choice through production rollout. That lifecycle approach is central to the product. Instead of treating an agent as a one-off script or a prompt that lives in a notebook, Foundry gives it a structured path from development to production.

Framework

Foundry agents run on Pydantic AI. The company avoided heavier agent frameworks such as CrewAI in favor of something simpler and more deterministic. Agents are defined as configuration, similar to infrastructure as code, and that configuration drives a common orchestrator with domain-specific code underneath it. The choice reflects a broader preference in production operations for systems that are predictable, testable, and easy to reason about. Deterministic behavior is especially important when an agent can trigger remediation, open tickets, or change network state.

Review and rollback

Customers commit agents to their own Git repository. Changes go through the customer's existing pull-request workflow. That means the same review, approval, and audit mechanisms used for application code can be applied to agent behavior. Before an agent reaches production, it is replayed against the customer's historical incident data and compared against the recorded outcome of each past event. A failed promotion rolls back in a single step.

Incident replay is a powerful trust mechanism. Historical incidents contain the messy signals that real networks produce: alerts, logs, topology changes, maintenance windows, provider notices, and operator actions. By replaying an agent against that record, teams can see whether the agent would have reached the same conclusion or taken a safe action. It also creates a feedback loop. When an agent fails a replay, the team can adjust configuration, add guardrails, or refine the workflow before production exposure.

Guardrails

Foundry applies two kinds of limits on agent behavior. The first caps cost and token usage, limiting how many model calls an agent can make before it is treated as broken. That is important because AI agents can loop, retry, or make unnecessary calls that increase cost without improving outcomes. Treating a runaway agent as broken rather than letting it continue is a practical operational control.

The second guardrail constrains what an agent is allowed to conclude. For example, if an agent is reporting an AWS outage, it cannot blame GCP. That kind of constraint prevents hallucination and reduces the risk of an agent making an illogical or unsupported diagnosis. In network operations, a wrong root cause can send engineers down the wrong path, delay resolution, or trigger unnecessary changes.

Autonomy

Full automation is the goal, but it is not expected to be the immediate reality. The company describes the target as completely non-human in the loop, while acknowledging that the transition will be gradual. Teams will likely move through stages: agent-assisted analysis, human-approved action, supervised automation, and eventually broader autonomous remediation. That asymptotic conversion reflects both technical limits and organizational trust. Even a capable agent needs to earn confidence across a range of incident types, change windows, and business-critical environments.

The workflows agents handle

Foundry agents are built around infrastructure support workflows. Handling one today means working through several sequential steps by hand, waiting for each one to finish, and then closing or opening a ticket depending on the outcome. That work is repetitive, time-sensitive, and often spans multiple teams and systems.

When something happens, an operator needs to check whether the underlying provider has maintenance going on. If it is not maintenance, the operator needs to file a ticket with the provider. When a cloud outage hits, an agent first has to determine whether the problem is real, for example whether a cloud provider such as AWS is actually down or the connection itself is fine. Troubleshooting an outage and determining what could be done is expected to be the most common workflow.

When an outage happens, an agent activates on its own and issues a plan of execution, followed by periodic status updates while the outage is ongoing. That plan might include checking provider status pages, validating network paths, correlating alarms, notifying stakeholders, and opening or updating tickets. The agent does not simply answer a question; it participates in an operational process. That process orientation is what separates an agent from a chatbot.

What's next

Agent interoperability is a major focus. The company expects Selector will not be the only platform agents run on, and the industry needs standard protocols, such as MCP and A2A, before agents built on different platforms can work with each other. Interoperability would allow specialized agents to cooperate across vendor boundaries, which is important in multi-cloud and multi-vendor networks. It would also reduce lock-in and make it easier for teams to compose agents from different sources.

Selector's main point of difference against established rivals is that customers can build their own agents instead of working from a fixed set the vendor provides. The platform is called Foundry for a specific reason: customers can create things on their own because their data is theirs and their workflows are theirs. A vendor can only guess what those workflows are, and prebuilt agents may not fit every organization's processes, compliance requirements, or network architecture.

That philosophy has trade-offs. Self-built agents require governance, testing, and skilled staff. But it also avoids artificial bounds on innovation. Teams can encode their own operational knowledge, escalate according to their own policies, and integrate with their own systems of record. In a field as varied as network operations, that flexibility may be more valuable than a one-size-fits-all catalog.

Another expansion is already underway into monitoring neoclouds, building on a cloud product Selector recently launched to compete with vendors such as Datadog. Existing data center monitoring tooling is not sufficient as that infrastructure gets built out, and the company expects to have customer results to share within six months. Neoclouds, which are specialized compute environments often used for AI workloads, have different traffic patterns, utilization profiles, and failure modes than traditional enterprise data centers. Monitoring them requires new data models and new operational workflows.


Source:Network World News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy