Thursday, October 8, 2026

The Experts Were Wrong About AI. Again.

 

Let’s explore how artificial intelligence has consistently outperformed the predictions made by experts in recent years. While many anticipated that significant AI milestones would take decades to achieve, models have frequently reached these benchmarks in just a year or two.

Key Areas of AI Progress (Underestimated)

  • Mathematical Problem Solving: In a 2025 survey, experts estimated a low probability for AI to contribute to solving a Millennium Prize problem by 2027. However, AI reached this benchmark in September 2026.
  • Maths Olympiad: Experts previously projected that AI would reach gold-level skills in the International Maths Olympiad by 2030, a milestone that was actually achieved in 2025.
  • Theorems and Publications: A 2023 survey of over 2,000 AI researchers predicted it would take 22 years for AI to prove theorems publishable in top mathematics journals. In reality, preprint servers are now being flooded with AI-authored papers, leading to new challenges regarding content responsibility and author comprehension.

 

Moving forward, physics and other theoretical disciplines may not require a new Einstein, but rather a systematic approach to utilizing AI to read and synthesize the vast existing body of scientific literature.

 

  • Programming Skills: AI's programming capabilities have far exceeded expert predictions, demonstrating rapid improvement in a short period.
    • LifeCode Bench Pro Performance: Experts previously predicted that the best AI score on this difficult coding test would rise to 14% by the end of 2026. However, by May 2026, AI had already achieved a score of 53.8%.
    • Task Completion Time: In April and May 2026, experts estimated that by the end of 2026, an AI would be able to reliably complete a coding job that takes a human expert about 3.4 hours. While the survey was still ongoing, a new AI model had already reached a time of 3 hours and 6 minutes

 

This rapid progress suggests that for digital tasks like coding, where institutional and physical hurdles are minimal, AI development is moving much faster than human analysts anticipated.

  • Cybersecurity: A significant acceleration in AI's cybersecurity capabilities, marking a clear departure from the conservative timelines previously projected by experts.
    • Rapid Capability Surge: In 2023, cybersecurity experts widely predicted that AI would not be capable of autonomously finding and exploiting system vulnerabilities for several years. However, AI reached this milestone in 2024, far ahead of schedule.
    • Autonomous Cyber Defense: While there is no single agreed-upon definition, the field is evolving toward agents that can move beyond simple threat detection to engage in active defense measures, such as system hardening and recovery. Ref: CSET (Center for Security and Emerging Technology)
    • Modern Cybersecurity Dynamics: Current cybersecurity now involves a race between AI-powered attacks and AI-powered defenses. While AI platforms are highly effective at detecting anomalies at machine speed, they often still require human oversight to distinguish between legitimate system maintenance/deployments and actual malicious activity.

 


Why Experts Were Wrong

  • Academic Bias: Many experts surveyed work in academia, an environment characterized by slow, methodical progress. They struggle to account for the velocity of market-driven economies, especially when those industries are supercharged by hundreds of billions of dollars in investment.
  • The Disconnect Between Bits and Atoms
    • Digital Success: Tasks that rely purely on software such as coding, mathematical proofs, and scientific writing face almost zero friction. These areas have seen rapid breakthroughs because they bypass physical constraints.
    • Physical Constraints: Conversely, domains requiring physical labor, massive infrastructure changes, or complex institutional shifts (like building power plants, mining, or changing corporate workflows) face significant, time-consuming hurdles that do not scale at the same speed as digital algorithms; institutional and logistical hurdles that will take decades.

 


Misplaced Predictions (Overestimated)

  • Job Automation: expert predictions regarding job automation have been largely incorrect, with the reality of workplace changes falling significantly short of past forecasts
    • Overestimated Impact: In 2020, the World Economic Forum projected that nearly 50% of work would be automated by 2025. A follow-up report in 2025 revealed the actual rate of automation was closer to 22%.
    • Specific Example: Roles that were expected to be automated early on, such as truckers and taxi drivers, have seen very little disruption so far, contrary to widespread predictions

The transition of physical, labor-intensive industries is much slower than the adoption of software-based AI tools. While digital tasks (like coding or writing) evolve rapidly, changing workflows in physical companies, building infrastructure, and updating supply chains face significant institutional and logistical hurdles that take decades to overcome

  • The 'Intelligence Explosion': The predictions of an imminent 'intelligence explosion' (by 2027) is proven unrealistic, given the immense physical infrastructure required to power it.
    • The Prediction: Forecast from Leopold Aschenbrenner, which posited an impending intelligence explosion by 2027. This theory suggested that AI would self-improve at an accelerating rate by building massive infrastructure, including hundreds of gigawatt-scale power plants and supercomputers, facilitated by robots designed by the AI itself.
    • The Critique: The scenario assumes an overly simplistic feedback loop where AI can easily command physical reality; ignoring the massive logistical, resource, and institutional hurdles involved in building power grids, mining, and large-scale industrial robotics.
    • The Reality Check: This "explosion" narrative contracts with the reality of how technical progress actually happens. While software tasks (coding, math) can indeed see near-zero friction progress, physical-world change (the 'atoms' side of the equation) requires decades of labor and coordination, making an instantaneous 'intelligence explosion' that transforms the physical world in a few short years highly unlikely.

 


Future Outlook: The Bifurcated AI Trajectory

I predict a bifurcated future for artificial intelligence, defined by the stark difference between digital-only tasks and physical-world implementation

  • Rapid Growth in Software Domains: Areas like mathematics, coding, and scientific research are expected to see exponential progress. Because these tasks occur within software environments and face minimal institutional or physical friction, we are entering a phase where AI will rapidly tackle long-standing theoretical problems perhaps even accelerating the discovery of new science.
  • The 'Physical' Bottleneck: In contrast, industries requiring massive infrastructure, energy, and material resources (such as power plant construction, large-scale mining, and industrial robotics) will likely evolve much more slowly. These transitions will take decades, tempering the likelihood of an overnight 'intelligence explosion' that drastically alters the physical world.
  • The Future of Human Contribution: Despite rapid AI gains, the speaker argues that progress will remain accessible. The near future may be characterized by a shift where AI tools enable almost anyone to function as a researcher, leading to a period of democratized scientific discovery, even as the ultimate societal impact whether it leads to total job displacement or a post-work utopia remains an open, debated question.


Monday, October 5, 2026

The Evolution of AI Agent Development: From Prompting to Autonomous Systems

 


The evolution of AI agents follows a clear architectural progression. Early applications focused almost exclusively on crafting better prompts. Today's enterprise-grade agents rely on multiple engineering layers that progressively increase intelligence, reliability, and autonomy.

Each new layer addresses specific limitations exposed by the previous one.

Rather than thinking of AI as "better models," it is more useful to think of it as an evolving software stack. Models remain the reasoning engine. The surrounding architecture increasingly determines real-world performance.

Layer 1: Prompt Engineering

Teaching the AI how to think
Prompt engineering represents the first generation of AI application development. Here, intelligence is largely encoded in the instructions given to the model.

Instead of simply asking a question, developers define:

  • Roles
  • Objectives
  • Constraints
  • Output formats
  • Reasoning style
  • Examples

Example:
Instead of: "Answer customer questions."

Better: "You are a senior customer support specialist for an enterprise SaaS company. Answer politely, cite product documentation, never speculate, and escalate billing questions."

The prompt essentially becomes the agent's temporary personality and operating manual.

Common techniques:

  • Role prompting
  • Few-shot examples
  • Chain-of-thought prompting
  • Step-by-step reasoning
  • Structured outputs (JSON, XML)
  • System prompts
  • Prompt templates

Strengths:

  • Extremely fast to build
  • Little infrastructure required
  • Excellent for isolated tasks
  • Works well for chatbots and assistants

Limitations:
The model only knows what exists inside its context window. It cannot:

  • Retrieve fresh information
  • Remember previous work
  • Use enterprise systems
  • Access databases
  • Interact with software
  • Perform multi-step workflows

Prompt engineering reaches diminishing returns because no matter how sophisticated the prompt becomes, the model remains isolated.

Layer 2: Context Engineering

Teaching the AI what it needs to know

Once developers realized prompts alone were insufficient, focus shifted toward supplying better information rather than better instructions.

This gave rise to Context Engineering.

Instead of embedding every possible fact into a prompt, the agent dynamically gathers the information it needs before reasoning.

Think of the context window as the agent's working memory. Context engineering determines what information enters that memory.

Typical sources include:

  • Enterprise documents
  • Databases
  • CRM systems
  • APIs
  • Web search
  • Code repositories
  • Vector databases
  • Previous conversations

This is the foundation of Retrieval-Augmented Generation (RAG).

Instead of hallucinating, the agent retrieves.

Example:

A procurement agent receives: "Compare these vendors."

Instead of relying on training data, it automatically gathers:

  • Existing contracts
  • Vendor performance history
  • Pricing
  • Compliance reports
  • Risk assessments
  • External market data

Only after collecting this information does reasoning begin.

Context Engineering Components:

Information Retrieval

  • Semantic search
  • Keyword search
  • Hybrid search

Memory Management

  • Short-term memory
  • Conversation history
  • Long-term memory

Tool Calling

  • SQL databases
  • ERP systems
  • CRM platforms
  • Internal APIs

Filtering
Only relevant information is added. Irrelevant context wastes tokens and reduces reasoning quality.

Why Context Engineering Matters:
Better models help. Better context helps far more.

Many enterprise failures blamed on "weak AI" are actually failures in context engineering.

Layer 3: Harness Engineering

Teaching the AI how to execute work

As organizations attempted increasingly complex workflows, another limitation appeared.

Large projects quickly exceeded available context windows.

Imagine asking an agent: "Build an entire ERP implementation plan."

Eventually:

  • Earlier decisions disappear
  • Objectives are forgotten
  • Inconsistencies emerge
  • Duplicated work appears

This is known as "context leakage."

The solution is not a larger model. The solution is an external orchestration layer.

This is Harness Engineering.

What is the Harness?
The harness surrounds the LLM. Instead of the model managing everything internally, the harness manages:

  • Tasks
  • Checkpoints
  • Execution state
  • Memory
  • Retries
  • Tool usage
  • Workflow progression

The model becomes only one component inside a much larger software system.

Example Workflow:
Suppose an engineering agent is asked: "Upgrade a 200-service microservice platform."

The harness may:

  1. Analyze repositories
  2. Create dependency graph
  3. Prioritize services
  4. Generate migration plan
  5. Execute updates
  6. Run unit tests
  7. Run integration tests
  8. Review failures
  9. Retry failures
  10. Produce deployment plan

Each task becomes an independent execution unit. The harness stores results externally. The model never has to remember everything.

Responsibilities of the Harness:

  • Task decomposition
  • State persistence
  • Checkpoint management
  • Retry logic
  • Tool integration
  • Execution monitoring
  • Result aggregation

Why the Harness Matters:
Longer workflows become possible. Complex multi-step processes can be orchestrated. The model focuses on reasoning, not bookkeeping.

Layer 4: Loop Engineering

Teaching the AI how to improve itself
Loop engineering represents the next evolution. Instead of relying on human prompts for every stage, this involves creating self-guided scaffolding outside the harness layer.

This allows agents to autonomously trigger tasks, perform maintenance, fix bugs, and verify their own work without constant human intervention.

What Loop Engineering Enables:

  • Autonomous verification
  • Self-correction
  • Iterative improvement
  • Bug detection and fixing
  • Continuous optimization
  • Long-duration autonomy

Example:
An agent is tasked with "Migrate this data pipeline."

In loop engineering, the agent:

  1. Executes the migration
  2. Automatically verifies results
  3. Detects any anomalies
  4. Investigates root causes
  5. Fixes identified issues
  6. Re-verifies
  7. Documents the process
  8. Escalates only if unresolved

The system operates over hours or days without human intervention.

Key Components:

  • Evaluation frameworks
  • Feedback loops
  • Error detection
  • Self-correction mechanisms
  • Autonomy guardrails
  • Human escalation paths

Why Loop Engineering Matters:
Agents can operate independently over long durations. This enables truly autonomous systems that improve over time rather than simply executing once and stopping.

The Progression of Autonomy

Layer Focus Model Role Human Role Complexity
Prompt Engineering Instructions Execute single query Write prompts Low
Context Engineering Information Reason with rich data Manage data sources Medium
Harness Engineering Orchestration Reason and decide Design workflows High
Loop Engineering Learning Reason, decide, improve Monitor boundaries Very High

What This Means for Builders

If you're building Layer 1 systems: Focus on prompt craft and few-shot examples. Understand that you've hit the ceiling on what prompting alone can achieve.

If you're building Layer 2 systems: Focus on information retrieval and context quality. Better context solves more problems than better prompts.

If you're building Layer 3 systems: Focus on orchestration, task decomposition, and external state management. The model is now a component inside a larger system.

If you're building Layer 4 systems: Focus on evaluation frameworks, self-correction mechanisms, and safety boundaries. How does the system know it succeeded? How does it catch and fix its own errors? How do you prevent drift?

The Competitive Advantage

Most organizations are still in Layers 1 and 2. They're writing prompts and managing context.

The organizations building Layers 3 and 4 right now will have a 2-3 year head start on complex, autonomous systems.

The architecture you build today determines what's possible tomorrow.

Friday, October 2, 2026

The Product Requirement Document Is Dying

 


You wrote a 47-page PRD. Half of it was ignored. Your designers misread the success metrics. Your stakeholders debated the screenshots. By the time everyone signed off, the market had moved.

PRDs aren't just struggling. They're becoming obsolete because the fundamental problem they solved has changed. Requirements matter more than ever. But the way we specify them has to shift.

Why PRDs Ever Existed

The PRD was born from constraint.

When a product manager worked in one building, architects in another, and engineers in a third, you needed a translation layer. Something to turn ideas into written specifications that people could interpret and execute.

The PRD answered six questions:

  • What are we building?
  • Why are we building it?
  • Who is it for?
  • How should it behave?
  • What does done look like?
  • What might go wrong?

It was a coordination mechanism. Imperfect, but necessary.

The Game Changes When Machines Execute

AI agents change the rules.

A coding agent can inspect a repository, analyze architecture, write implementation plans, generate code, run tests, and iterate. An agentic product system can research markets, analyze customer feedback, generate designs, create plans, call APIs, and coordinate multiple specialized agents.

The question shifts entirely. A PRD works fine for coordinating humans. But an intelligent system needs something different: it must interpret intent, execute within boundaries, and continuously demonstrate that an outcome is being achieved. That requires a completely different artifact.

Why Static PRDs Fail AI Agents

When an agent encounters a traditional PRD, it faces five fundamental problems:

1. Intent without boundaries. PRDs describe aspirations ("make onboarding frictionless") but not operational boundaries ("onboarding can take 5-7 minutes, not longer"). An agent needs to know what counts as success, what states are valid, what must never happen, which tradeoffs are acceptable.

The next problem is worse.

2. Ambiguity at scale. A human engineer who reads "users can easily reset their password" can ask clarifying questions and infer conventions from the codebase. An agent chooses a plausible interpretation and implements it at scale. If the interpretation was wrong, you discover it in production.

3. Missing context. A PRD sits in a document system while the code, APIs, data models, design system, and deployment rules live elsewhere. An agent cannot safely plan from product intent alone. It needs the relevant architecture, existing conventions, known technical debt, security policies, and prior decisions.

This compounds quickly.

4. Linearity. The classic PRD implies requirements → design → code → testing. Agents work differently. They propose alternatives, build prototypes to resolve uncertainty, discover conflicts with existing APIs, generate tests while writing code, ask other agents to critique their plans, and revise their task graphs based on results. A static document can't keep pace with this.

5. Missing accountability. When a human team implements a feature, responsibility flows through recognizable roles. When agents plan and execute, the question becomes: who authorized the agent to make this decision, on what evidence, under which policy? A PRD rarely records this. An agent operating in production needs an audit trail, not a 30-page narrative.


Left shows traditional PRD as a thick narrative document addressed to humans. Right shows modern stack as interconnected layers (intent, behavioral, contract, evaluation, execution) addressed to AI agents. Notice the shift from prose to precision.


What Replaces the PRD

The answer is not "a better PRD 2.0."

The answer is a living product execution system composed of:

  • A human-readable statement of intent
  • A machine-readable behavioral specification
  • Explicit context and policy layers for agents
  • Executable acceptance tests and evaluation criteria
  • A dependency-aware task graph
  • Traceable evidence from implementation through production
  • Clear approval gates for decisions that humans must make

Think of it as moving from:

"Here's what we want you to build."

To:

"Here's the problem, desired outcome, context, constraints, policies, evidence thresholds, and definition of success. Determine the best way to achieve it, operate within these boundaries, and continuously demonstrate that the outcome is being achieved."

What This Looks Like in Practice

Instead of a 25-page narrative, product teams maintain five layers:

1. Product Intent Layer (human-facing, narrative)

This is how you communicate why. It replaces the opening sections of traditional PRDs:

  • The customer or business problem
  • Why it matters now
  • Target users
  • Desired outcome
  • Strategic constraints
  • Non-goals
  • Ethical or regulatory boundaries

This layer stays readable. It's for strategy alignment, not machine execution.

2. Behavioral Specification (machine-readable)

This is how you communicate what. It's precise enough that an agent can plan from it:

  • Actors and permissions
  • Inputs and outputs
  • User-visible states and transitions
  • Business rules
  • Error behavior and edge cases
  • Performance and reliability requirements
  • Explicit non-goals

3. Agent Operating Contract

If an agent is executing your product strategy, it needs to know its boundaries. This layer replaces handoff meetings:

  • Which repositories the agent can access
  • Which tools it can call
  • Which files it can modify
  • Coding and architectural conventions
  • Security and data-handling rules
  • When it must ask for approval
  • Which sources are authoritative

4. Evaluation Contract

How do you know the agent succeeded? This replaces post-launch review meetings:

  • Representative test cases
  • Adversarial cases
  • Golden outputs or acceptable ranges
  • Quality rubrics
  • Safety and refusal criteria
  • Latency and cost budgets
  • Production feedback signals

5. Execution Graph

This is the work breakdown, generated dynamically:

  • Epics, capabilities, tasks
  • Technical and design dependencies
  • Data migrations
  • Rollout stages
  • Approval gates

The key insight: you're no longer trying to fit everything into one document. Each layer has one job.


The complete system showing Product Intent, Behavioral Specification, Agent Operating Contract, Evaluation Contract, and Execution Graph as interconnected layers. Agents interact with the system on the right, feedback loops flow upward. This is the operational structure that replaces a static 30-page document.


The Role Changes

The product manager doesn't disappear. The role evolves.

Instead of coordinating execution, PMs become intent architects and control-plane designers.

Your week changes from:

  • Writing tickets
  • Grooming backlogs
  • Clarifying requirements
  • Attending status meetings

To:

  • Defining strategy and the outcomes that matter
  • Analyzing what the data actually says
  • Designing and running experiments
  • Setting constraints and approval gates
  • Watching agent behavior in production
  • Reviewing decisions that affect revenue, risk, or customers
  • Predicting what might break

The question changes from:

"How do I describe this feature in enough detail that engineering understands?"

To:

"What should the system optimize for, under what constraints, using what evidence, with what level of autonomy?"

Start Now, Not Later

Here's the uncomfortable truth: you don't need to wait for agents to start this shift. Your current human teams would move faster if you stopped writing feature narratives and started writing intent systems. The PRD is holding you back whether you're using agents or not.

This transition doesn't require abandoning PRDs tomorrow. It requires changing what they contain right now.

Stop optimizing for: "Can engineering implement this requirement?"

Start optimizing for: "Can an intelligent system understand the objective, constraints, evidence, and success criteria?"

Start adding:

  • Explicit outcomes (what measurable result?)
  • Explicit constraints (what must not happen?)
  • Explicit decision rights (what can an agent decide?)
  • Explicit evaluation criteria (how will we judge?)
  • Explicit context (what should an agent consult?)
  • Explicit escalation rules (when must humans intervene?)
  • Explicit telemetry (what evidence proves it's working?)

This gradually transforms a PRD into an agent-ready specification.

The Final Shift

The fundamental change is this: execution is becoming autonomous.

When humans do the work, you need detailed instructions. When agents do the work, you need clear intent, relevant context, hard constraints, evaluation mechanisms, governance, and feedback loops. The PRD was built for a world of human execution. We need new artifacts for a world of autonomous execution.

The product manager's job shifts upward in the abstraction stack.

You stop saying "build this" and start saying "achieve this." You stop describing implementations and start defining boundaries. You stop writing acceptance tests and start designing evaluation mechanisms.

The PRD's real value wasn't the document itself. It was reducing ambiguity between teams. Now that teams include machines, ambiguity creates production risk. An agent will confidently execute an ambiguous requirement in ways that break your product.

The product manager's real job becomes designing the system in which intelligent agents make good decisions.

What this means for you next week:

Pick one upcoming feature. Instead of writing a 15-page PRD, write:

  1. A 300-word problem statement (human-readable intent)
  2. Five concrete success criteria (behavioral spec)
  3. A one-page "what must never happen" list (boundaries)
  4. How you'll know it's working in production (evaluation)

That's your starter template. No agents required. Your current team will ship faster.

The successor to the PRD isn't just a better document. It's a system of durable intent, clear boundaries, and continuous evaluation. Start building it now not when you have agents, but when you have the next feature to ship.

Sunday, September 27, 2026

Jev: The Model Built for Decisions, Not Text

 


AI system makes a decision. It returns text. Your code has to parse that text, interpret what it means, and act on an educated guess about intent.

A routing system sends customer inquiries to the wrong department. A content filter misses harmful content. A risk scorer gives you a number wrapped in prose your engineers have to extract.

Most AI deployments are fighting this friction every day.

The bottleneck isn't the model's intelligence. It's the layer between the model's output and your code. Every time you insert that layer, you introduce risk: hallucination, parsing errors, ambiguity.

TypeSafe's Jev changes this fundamentally.

The Invisible Tax Your Engineering Team Pays Every Day

Every production system using LLMs to make decisions faces the same architectural problem:

LLMs output text. Code requires structure.

This creates a three-step process that shouldn't exist:

  1. Model generates text (often verbose, sometimes ambiguous)
  2. Your system parses the text to extract the decision
  3. Code uses the extracted value and hopes the interpretation was correct

Each step adds latency, cost, and fragility. The model can hallucinate. The parser can misinterpret. The extracted value might be ambiguous.

Most teams respond by adding guardrails, validation layers, retry logic, and fallback mechanisms. You're spending significant engineering resources translating between model outputs and code inputs.

This is why AI systems that work in demos fail in production. The demo glosses over parsing. Production systems can't afford to.

What Actually Is Jev?

Jev isn't a text generator. It acts as a generalized classifier to make fast, structured decisions that can be consumed in a software. It's a decision engine that returns typed values.

TypeSafe describes it as a System One model borrowing from Daniel Kahneman's research on fast, intuitive decision-making. But in engineering terms, the distinction is functional: Jev is optimized for bounded decisions while traditional LLMs remain optimized for open-ended text generation.

Instead of generating prose, Jev evaluates questions and returns one of three structured outputs:

1. Choice: Select from predefined options with confidence scores

  • Example: Route this support ticket to Sales, Support, or Engineering?
  • Output: {selected: "Support", confidence: 0.94}

2. Score: Rate something against a rubric on a numeric scale

  • Example: How risky is this transaction? (1-10 scale)
  • Output: {score: 7, confidence: 0.87}

3. Noul: Determine if a statement is true or false (probability 0–1)

  • Example: Does this content violate our policy?
  • Output: {truth_value: 0.92} (92% confidence it violates)

No parsing required. No text extraction. No hallucination risk. Structured data in, structured answers out.

The Architecture Breakthrough

Here's where Jev separates from traditional approaches: all questions execute in parallel within a single request.

Imagine you need to evaluate a customer for five different criteria: risk score, segment, churn likelihood, lifetime value band, and next-best-product recommendation.

With a traditional LLM, you either:

  • Make five separate API calls (slow and expensive)
  • Chain them sequentially (even slower, and each step creates compounding hallucination risk)
  • Prompt-engineer a single complex query that tries to answer all five at once (fragile, easily confused)

With Jev, you ask all five questions simultaneously. The model answers each independently with explicit confidence scores on every decision.


The impact? Fewer model calls. Lower latency. Lower API costs. More reliable decisions.

Why This Matters: The Business Case

If you're building systems that use AI to make decisions not to generate creative content, but to route, filter, rank, or branch logic; Jev changes your cost structure and reliability profile.

Cost: Parallel evaluation means fewer model calls. In high-volume decision systems (routing thousands of support tickets, scoring millions of transactions, filtering millions of pieces of content), this multiplies across your infrastructure costs. Jev claims 80x+ cost reduction.

Latency: Traditional approaches add overhead at every decision point. Jev's structured outputs remove that friction. In real-time workflows, milliseconds matter. Jev offers significant advantages in speed (sub-500ms).

Reliability: When the model returns text, schema failures happen; formatted wrong, missing quotes, ambiguous output. Jev returns typed values, eliminating that entire class of failure. A confidence score on a routing decision is unambiguous: 0.92 means 92% confidence, not prose that your engineer has to interpret.

Developer Productivity: Teams spend less time building validation layers, retry logic, and guardrails. The model outputs what your code expects, not text your engineers have to sanitize and parse.

Auditability: Confidence scores provide explicit traceability. You can defend decisions: "Routed to Support with 94% confidence based on topic classification." For regulated workflows, this matters.

Where Jev Wins (And Where It Doesn't)

Jev excels in specific types of problems:

  • Routing: Which queue, department, or specialist?
  • Filtering: Does this meet the criteria?
  • Ranking: How does this score across multiple dimensions?
  • Branching logic: Which path should this workflow take?
  • Risk assessment: How risky is this?
  • Qualification: Does this lead meet the bar?

If your use case is "I need the model to generate creative marketing copy, write code, or reason through a complex problem," Jev isn't the tool. Use GPT, Claude, or another generative model.

If your use case is "I need the model to make a fast, reliable decision my code can act on immediately," Jev changes the economics and architecture of how you build systems.

The Architectural Shift This Represents

Industry treated LLMs as general-purpose tools and forced them into decision-making roles. We've built layers of prompt engineering, parsing logic, and validation to make it work.

Jev represents a different philosophy: purpose-built architecture for the specific job it's built for.

Think of a modern AI system in two complementary layers:

  • Generative layer: writes, explains, summarizes, plans, reasons, and interacts with people
  • Decision layer: classifies, scores, gates, routes, filters, and determines whether the next action should occur

You don't need both layers for every workflow. But executives should stop assuming one general-purpose model can efficiently perform every cognitive function.

A useful enterprise pattern: Business event → state/context → Jev decision → software rule → action, escalation, or generative model.

The critical principle: the model makes the judgment; your application code owns the consequences.

The Question for Engineering Team

If you're evaluating AI infrastructure, ask this: "Are we spending engineer-years building validation layers around text generation, or are we using architecture designed for structured decisions?"

Jev is one answer. Not the only answer, but it signals a shift in how the industry thinks about integrating AI into software.

The move from "How do I get a text-generating model to make decisions?" to "How do I build systems that make structured decisions?" is subtle. But it's profound.

Wednesday, September 23, 2026

Why Most AI Products Are Just Chatbots Wearing Makeup

 


You've seen this a hundred times.

A polished demo. Natural language input. A confident pitch about "AI-powered transformation." But strip away the interface and branding, and you're left with the same interaction pattern: type a message, get a generated response. Chatbot. Different logo, same chat window.

The problem isn't that these products are useless. It's that most teams, and most PMs, can't tell the difference between a product that uses an LLM and a product that is an LLM wrapper with a login page. That confusion costs money.

The Chatbot Wrapper Epidemic

Here's why this keeps happening:

Speed beats substance. Building a chat interface takes days. Building a real product takes quarters. When leadership wants "an AI strategy" by end-of-quarter, a text box is the fastest thing to ship. Investors reward velocity. The market rewards demos. So teams ship chatbots.

Chat hides the hard decisions. A blank text box avoids the product thinking: Which problem? Which workflow? What counts as done? When the interface is "ask anything," nobody has to answer those questions. The customer does that work instead.

Foundation models make it easy. With a single API call, any team can make something that sounds intelligent. That lowers the barrier to shipping but also lowers differentiation. If your value proposition is just "we wrote a good prompt," it's not defensible.

Here's the uncomfortable part: most teams know this. They know they're shipping makeup. They ship it anyway because the market doesn't punish it fast enough. Investors reward the demo. Early users treat it like a feature. By the time competitors ship something real, your team is already staffed up and your roadmap is locked in. But that's also where the risk lives in the gap between "this shipped" and "this matters."

The Test: Remove the Chat UI. What's Left?

Here's the first question every PM should ask about an "AI product":

If you remove the conversational interface, does the core capability still work?

If the answer is yes; if the product is still valuable; you might have something real. A routing engine, a classification system, a workflow automation. Something that solves a problem differently.

If the answer is no; if it falls apart without the chat; then the chat was the product. Everything else was theater.

Most products in the market fail this test.

Five Levels of AI Product Maturity

Not every "AI product" is created equal. Think of them on a spectrum:

Level 1: AI Interface
The existing product unchanged. AI provides a new way to interact with it.
Example: "Ask our CRM anything." You type questions; AI answers. The underlying workflow stays manual. You still read, decide, copy, paste, execute.

Level 2: AI-Assisted Workflow
AI begins participating in the work, but humans remain the primary orchestrator.
Example: Support ticket arrives → AI summarizes → AI proposes diagnosis → human approves → system updates ticket. You accelerated individual steps, not the workflow.

Level 3: AI-Orchestrated Workflow
The system coordinates the work, not the user.
Example: Support ticket arrives → system classifies → retrieves history → determines resolution → updates CRM → escalates exceptions. You provide an objective; the system executes a process.

Level 4: AI-Native Product
AI is core to the product architecture. The product couldn't exist in this form without it.
Example: A system that continuously observes data, interprets conditions, generates hypotheses, makes decisions, executes actions, evaluates outcomes, and learns from feedback.

Level 5: Adaptive System
The product learns from outcomes and adapts behavior based on changing inputs, feedback, and policies.
Example: A financial system that not only processes transactions but improves its routing logic based on market conditions and outcome patterns.

A product doesn't need Level 5 to be valuable. A narrow Level 3 can create enormous impact if it reliably solves a painful, expensive, frequent problem. The point is: is there actual workflow change, or just a new interface?


Five Questions That Separate Substance From Makeup

1. Does it remove work or create another step?

A superficial product gives users one more place to ask questions. A strong product eliminates steps in the end-to-end job.

  • Weak: A meeting-notes chatbot that transcribes and summarizes a call.
  • Strong: The system automatically links the meeting to the account, identifies commitments, assigns follow-ups, updates the CRM, drafts customer communications, tracks completion, surfaces unresolved risk.

2. Does the system understand context?

Generic AI produces generic assistance. Real products work with the right context at the right moment.

  • Weak: A chatbot that answers questions with public knowledge.
  • Strong: A system that knows the employee's role, department, payroll jurisdiction, benefit eligibility, prior cases, and relevant policy version. It produces a decision that is actually useful.

3. Can it safely execute?

Text generation is not execution. I watched a support team adopt an AI solution that drafted perfect responses but never updated the customer record. Six months in, they were maintaining two systems: one for what the AI said, one for what actually happened. Products become materially more valuable when they connect to systems of record, use tools, and produce verified changes; not just suggestions.

  • Weak: "Here's what you should do."
  • Strong: It updates the CRM, schedules the technician, prepares the claim, flags fraud, initiates the approval workflow. Safely.

4. Is performance measured on outcomes?

Many AI products are judged by impressive examples, not systematic performance. That's a red flag. Ask the vendor to share their worst-case scenario, not their best. Ask what happens in the long tail of edge cases. If they can only show you the highlight reel, you're looking at a demo, not a system.

Real products have metrics: Resolution time down 40%. First-pass approval rate up 65%. Error rate below 2%. These are task-specific measures.

5. Does each interaction make the product better?

A generic chatbot starts every interaction with roughly the same capabilities. A true product accumulates advantage.

  • Weak: You ask the same question next month, get the same quality answer.
  • Strong: User feedback, edits, approvals, and outcome labels flow into operational learning. The system improves from your usage of it.

What to Build (Or Buy)

If you're shipping an AI product, don't start with the chat box. Start here:

What painful workflow would disappear if this problem was solved?

Name it. Measure it today. Then ask: does AI solve this differently? If yes, does AI fundamentally redesign how the work gets done?

If the answer is yes, build around that capability. Design the interface for the outcome, not for the model. Maybe it's a dashboard. Maybe it's automated workflows that don't require user input at all. Maybe it's a hybrid: chat for exceptions, structured forms for routine work.

Then measure whether it moved the needle. Not "our users can now chat with the system." But "we reduced processing time from 5 days to 8 hours" or "we eliminated 60% of manual categorization" or "customers resolved issues without escalation 3x more often."

The Uncomfortable Truth

Most AI products shipping today are feature releases, not products. They're incremental UI improvements wrapped in the promise of AI. They tap into real demand and real FOMO. They attract funding and headlines.

But they're not solving fundamental problems differently.

The products that will matter in two years are the ones that can answer honestly: "Without AI, could this problem be solved at all?"

If the answer is yes; if it could be solved, just slower or more expensively; you're competing on efficiency, not innovation.

If the answer is no; if this problem didn't exist before AI, or couldn't be solved before; you're building something real.

Most of what's being marketed as AI products right now? Chat layer on top of a thing. That thing would work fine without the chat.

Smart PMs are asking harder questions before they ship. They're asking: "What would we do differently if we couldn't use a chat interface?"

That question will separate the real products from the dressed-up ones.