| A View from the Watchtower |
Design, manufacturing and the data center are each adopting AI agents. The value shows up in the connections between them, and today almost no one is managing those connections.
At SEMICON West, a lot of the conversation will be about the true value and cost of agentic AI in SoC design and manufacturing. It's the right question. From the watchtower, the answer is becoming clear: agentic AI pays off most not inside any one tool, but in the spaces between design, fab and deployment, where the most expensive problems start.
Every complex program runs on human glue: the architect who remembers why a margin was set, the product engineer who notices a test result doesn't match the model, the program manager who realizes two email threads are about the same problem. For decades that glue held. It was never designed; it was people being good at their jobs. It's failing now because the systems it holds together have outgrown what any person can track.
One change, from design to data hall
Take one change and follow it.
A model team revises an AI workload. Utilization rises, and the accelerator will run hotter for longer. The chip's thermal envelope moves, and on paper it's just a revised number in a spec.
In design, that number reopens power and timing closure. It changes which blocks need attention and which margins still hold.
In manufacturing, it reaches further than most design teams see. Test limits and burn-in conditions need revisiting. Binning strategy shifts, because the parts that meet the new envelope may be a different slice of the distribution. The package may need a different thermal solution, which can mean a different substrate, assembly flow or OSAT qualification.
In the data center, heat per rack goes up. The rack may cross the line where air cooling stops being enough and direct liquid cooling becomes mandatory. That changes coolant distribution capacity, the facility water loop and the mechanical design of the data hall. Those are long-lead, capital-intensive decisions, often locked in many months before first silicon ships.
|
The same change, caught at step 2 or at step 6. The difference is the cost of change.
|
The ripple runs the other way too. Early yield and test data might show a bin running hotter than the model predicted. That's a fab signal with consequences for the design team's next revision and for the facility planner's cooling budget. Neither of them is on the distribution list.
In a data hall, "set concrete" is not a metaphor: it's poured, it has piping in it, and changing it is very expensive. Digging it up costs far more than getting the pour right at the start. The same holds for a mask set, a qualified package or a locked test program.
Nobody in that chain did anything wrong. Each team was correct inside its own boundary. What failed was the space between them.
Every silo is getting its own AI
Design tools are gaining agents for simulation, placement and verification (EDA Is Entering Its Third Era). Fabs are applying AI to process control and yield learning (Photonics Is Driving a New Approach to Yield Analytics). Data center operators are using it to plan power and cooling. These are real gains.
But every one of them stays inside its own boundary. A design agent doesn't know the test program changed. A yield model doesn't know the workload moved. A facility planner's model doesn't know either. The problems that break schedules and budgets sit at the handoffs. That's why we argued the engineering stack itself has to be re-architected (Re-Architecting the Engineering Stack), and why so many issues still surface late (The Validation Crisis in AI Silicon; Why Are We Still Finding So Many Bugs So Late?).
Making each silo smarter doesn't connect the silos. It just gets each one to the wrong answer faster.
What an orchestration layer actually needs
An intelligent orchestration layer that spans design, manufacturing and deployment is the obvious next step. It would be made of agents that watch across domains, trace a change to everything it touches, and raise the issue early. But agents are only as good as the context they're given. Three foundations have to be in place first.
A living, end-to-end requirements thread
Intent has to be captured as connected, versioned requirements, from the workload through design, fab, test, package, system and facility. It can't live in documents that go stale the day they're approved. The thread has to survive handoffs between companies too, because the chip designer, the foundry, the OSAT, the ODM and the data center operator are rarely the same organization (The Spreadsheet Was Never the Plan; From Intent to Yield; Tracking Requirements Across Company Boundaries; Functional Safety Was the Front Door, Not the House).
A complete, dynamically linked bill of materials
Every IP block, derivative, die, process option, package variant and board-level component has to be linked to the requirements it satisfies. When something changes, you should be able to ask "what does this touch?" and get an answer in seconds, not after a two-week spreadsheet exercise. That's the blast-radius problem, and generic PLM wasn't built for it (The IP Blast Radius Problem; Silicon IP Operations: Your PLM Can't Do It, and Neither Can Jira; Why the Simple BOM Is Now Such a Big Deal).
Design intent that can be re-run
When a requirement moves, someone has to ask what the design would do now. If the only answer comes from a full RTL respin, the question won't get asked in time. Higher-abstraction, executable models change that economics (AI Will Reshape Chip Design, But Not Through RTL Generation Alone; When the First HLS Successes Stop Scaling).
|
Without these, an orchestration layer is guessing. With them, it can reason.
The right human, on the right problem, at the right time
With that context in place, the orchestration layer does something no single tool can. It notices that a requirement moved, or that a yield signal doesn't match the model. It walks the thread to find everything upstream and downstream. It traces the linked BOM to see which dies, packages, test programs and suppliers are affected, and re-runs the models that can be re-run quickly. Then it takes the result, with the evidence attached, to the one person who needs to see it: the design lead, the product engineer, the packaging lead or the facilities planner.
It doesn't replace judgment. It moves judgment earlier. Depending on the change, that's hours, days, weeks or months ahead of when conventional methods would have caught it, if they caught it at all. That's where the value of agentic AI shows up: in the difference between adjusting a drawing and breaking up concrete.
Beyond EDA 3.0
At AiT we've called this EDA 3.0: orchestrating the semiconductor lifecycle from intent to yield rather than speeding up individual steps. Our look at the memory wall in AI inference silicon (The Memory Wall Is Reshaping AI Inference Chip Design) pointed the same way.
The ripple above shows the idea doesn't stop at the chip, or even at the fab. Design, manufacturing and the data center are one industrial system. The pattern gets far bigger when it's paired with an industry leader that is bringing AI into every stage of the industrial process. One of the largest technology companies has already moved an agentic R&D platform into general availability, built to coordinate data analysis, hypothesis generation, experimentation and knowledge management through specialized AI agents. It has named semiconductors and advanced manufacturing among its target domains.
See it working, the same evening
This is no longer a whiteboard idea. Working demonstrations are being built on real designs, with real requirement chains and real change requests flowing through them.
If you're at the IEEE panel on the true value and cost of agentic AI in SoC design and manufacturing, Tuesday, October 13 at 12:15 PM in Moscone North Hall, Room 20AJ, you'll hear the question framed well. That evening we're showing one answer, live, at a small invitation-only gathering near Moscone. Request an invitation and we'll save you a spot.
Simon Bennett, CSO and Co-Founder, AiT
|
Signals worth watching
The building blocks referenced above, for readers who want to put the pieces together.
Living requirements: Jama Connect
Jama's semiconductor solution is built around Live Traceability across tools and engineering disciplines, which Jama positions as the answer to rework, respins, cost increases and delays. Jama Connect for Semiconductors announcement
Linked BOM and IP lifecycle: Glide Systems
Glide's SysLM platform pairs PDM/PLM with OSLC-based integrations to tools such as GitLab and Jira, aimed at implementing the digital thread. Glide SysLM general availability
Executable design intent: Rise Design Automation
At DAC 2026, Rise launched IPCreate and IPTransform, AI-assisted platforms that automate parts of the flow between specifications, HLS and RTL; Rise says its HLS models simulate about 30 times faster than equivalent RTL. Launch coverage and SemiWiki CEO interview
Agentic orchestration: Microsoft Discovery
Microsoft describes Discovery as an extensible platform combining agentic orchestration, advanced reasoning, a graph-based knowledge foundation and high-performance computing for R&D workflows. Microsoft Discovery documentation and Moor Insights on the Build 2026 launch
|
|