Skip to the essay
AMUI Design
AMUI Journal AI & product development8 min read · Interactive

If AI makes coding 10× faster, why doesn’t the whole company become 10× faster?

Suppose a feature takes seventeen days from “we should build this” to “it actually works.” Ten of those days are coding.

What changes if only those ten days collapse?

An AMUI interactive interpretation inspired by Andrew Ng. The 10× speedup and project numbers are hypothetical, not a measured productivity claim.

The answer is not 1.7 days. Because the other seven days do not move ↓
Move only the coding time. The rest of the work stays exactly where it was. This simplified example assumes sequential work and holds the other stages fixed; real teams may overlap tasks or improve several stages at once.
17.0 days total
overall speedup1.0×
Understand
Decide
Build
Check
3 days2 days10 days2 days
Understand3 days
Decide2 days
Build10.0 days
Check2 days
10.0d

Even if coding took no time at all, understanding, deciding and checking would still take seven days.

7-day floor
At ten days of coding, most of the calendar is still engineering.

Drag the build step down to one day and the whole project falls from seventeen days to eight.

Ten times faster at coding. Just over twice as fast overall.

And that changes the expensive mistake.

Yesterday, a weak idea could waste ten engineering days. Tomorrow, AI can help you build the same weak idea before lunch.

If building gets cheap, choosing badly gets expensive.

Imagine our team runs an online shop.

Customers keep abandoning checkout. The team can ship one change tomorrow.
1

Customers repeatedly leave when shipping cost appears for the first time at the end.

2

Support messages ask why the final price is higher than the cart total they saw earlier.

3

No current evidence shows customers asking for product recommendations during checkout.

AI recommendations
0
Shipping cost earlier
0
AI made both ideas cheaper to build. It did nothing to make them equally worth building.

Suppose the team chooses well.

They decide to fix another expensive problem: refund requests that bounce between customers and support staff because nobody has the right information at the right moment.

The first version of the assistant is ready almost immediately.

A customer asks for a refund. The assistant answers perfectly.

There is only one problem.

It made the refund policy up.

Fix that failure and the next one appears. Each fix gives the assistant one more ability. And one more way to fail.
Customer: “Can I get a refund for order 482?”
Customer
AI assistantMODEL
Answer / action
RETRIEVAL
TOOL
WORKFLOW
AGENT
Customer asks about order 482
↓
Refund policy
Order 482
↓
AI assistant MODEL
WORKFLOW eligibility → order → confirmation → refund
ASKclarify
LOOK UPinspect
ACTor escalate
↓
Answer / action
“Absolutely. Your purchase is covered by our 30-day refund policy.”
The answer sounds reasonable. The company does not actually have a 30-day refund policy.

The last version is different.

It can look things up. It can inspect real business state. It can follow a fixed procedure. And when the procedure stops fitting the case, it can choose what to do next.

That is the point where the word agent becomes useful.

Now the system works in the demo.

One customer asks for a refund. The assistant checks the policy, checks the order, follows the procedure, and stops exactly where it should.

Would you deploy it?
One routine refund worked. We still do not know where the assistant breaks.
CASES SEEN1 / 30
ILLUSTRATIVE RESULTS · PASS RATE100%
WHERE IT BREAKSNot enough cases yet

One green case proves the assistant can work. It does not tell us where it stops working.

Once the rest of the cases appear, the interesting result is not just the pass rate.

The failures cluster.

Conflicting information causes trouble. Missing data makes the assistant overconfident. Adversarial inputs push it away from the intended process. Strange order states break assumptions the happy path never exposed.

A demo asks, “Can this work?” Evaluation asks, “When does it stop working?”

And once the system can act, there is one more trap.

The assistant issues a refund and reports “success.” The request was accepted. The job was created. From inside the agent, the task looks finished.
What the agent can see
Refund request sent
API: accepted
Refund job created
SUCCESS
What happened in the world
Refund job created
Customer ledger unchanged
Customer still has no refund
Refund request sent ✓
API accepted ✓
Refund job created ✓
SUCCESS
Refund job created ✓
Customer ledger unchanged ×
Customer still has no refund ×
The agent wasn’t lying.
It checked the wrong definition of success.
The action succeeded. The outcome did not.

The agent checked the action it took.

Verification checks the state the world ended up in.

That is the deeper answer to the original question.

AI can make one part of building software astonishingly cheap.

But useful software was never just the code.

If coding becomes cheap, what becomes valuable next?

Choosing what deserves to exist. Understanding the system well enough to see how it can fail. And proving that the result is true outside the model that produced it.

DEPTH → understand failure
BUSINESS FOCUS → choose the right problem
DELIVERY → make the result hold up
Source note

From the lecture: rapid AI progress, faster AI-assisted software development, growing importance of product judgment, depth, business focus, delivery, agents and responsible deployment. Our teaching examples: the 17-day project, online-shop decision, refund assistant, evaluation matrix and refund-verification incident are invented examples used to expose those ideas. They are not quotations from Andrew Ng or Laurence Moroney.

Putting it into practice

What this means for your next product.

Before accelerating a build, choose one user task and define a result you can observe. Map the information the system needs, the actions it may take, and the cases that require human review. Faster implementation earns its value when the whole workflow works.

A hypothetical AMUI project: imagine an online shop considering a refund-support assistant. We would evaluate resolution accuracy: did it apply the current policy to the actual order and reach the correct outcome? We would also evaluate escalation behavior: did it hand uncertain, conflicting, or sensitive cases to a person before taking action, with enough context for that person to help? Compare those results with the shop’s existing support process. Faster coding makes the prototype easier to build; evidence about correct resolutions and appropriate escalation tells us whether it deserves to ship. This is an invented example, not a client case study or a measured result.

Explore web application development ↗