Tag: ai development

  • **What Will Remain Human in Software Architecture? A Focus Group Report: The Evidence and Implications**

    **What Will Remain Human in Software Architecture? A Focus Group Report: The Evidence and Implications**

    Header image: ‘Sooner or later one has to take sides to remain human’, The Quiet American, buildings blur, Seattle, Washington, USA by Wonderlane, CC BY 2.0, via flickr via Openverse — cropped to 16:9 and colour-adjusted.

    Key takeaways

    • Humans retain decision-making, accountability, and guardrail authoring in AI-assisted architecture
    • AI-generated code shows 75% more logic errors and 2.74× more XSS vulnerabilities
    • Criticality (uncertainty + cost of change) defines the need for human oversight

    AI can generate code. It can suggest patterns. It can even optimise performance. But when the stakes rise—when uncertainty meets high cost of change—humans still decide. And when things break, humans still answer. That’s not speculation. It’s the blunt consensus from a EuroPLoP 2026 focus group of 22 industry and academic practitioners.

    No caveats. No futurist hedging. Just a clear line in the sand: architectural decision-making, accountability, and guardrail authoring stay human. The data backs it up. AI-generated code carries 75% more logic errors, 2.74× more XSS vulnerabilities, and 30% higher change failure rates. Pull requests are 18% larger—evidence of less efficient, more fragmented solutions. These aren’t edge cases. They’re systematic limitations.

    The group didn’t just observe the numbers. They diagnosed why. Humans remain irreplaceable in three domains: judgement under uncertainty, ethical accountability, and designing the rules that govern AI itself. The emerging discipline of harness engineering isn’t about replacing architects. It’s about equipping them to govern the tools now assisting them.


    Harness Engineering: The Discipline of Governing AI-Assisted Architecture

    Harness engineering isn’t a buzzword. It’s the necessary response to AI’s limitations. The EuroPLoP 2026 focus group defined it as three layers: validation mechanisms, knowledge layers, and governance frameworks.

    Validation mechanisms—automated testing pipelines, compliance checks, static analysis—have shifted purpose. They no longer catch human errors. They catch AI-generated ones: logic flaws, security vulnerabilities, architectural constraint violations. But these mechanisms must be designed. Someone must define failure thresholds, review triggers, and acceptable risk levels. That someone remains human.

    Knowledge layers store company-specific standards, patterns, and domain models. AI can suggest a microservice design. It can’t decide whether that design aligns with a company’s tolerance for operational complexity or its long-term roadmap. That knowledge must be codified—by architects, product owners, even legal teams—and fed into the system. The focus group emphasised this as where organisational culture lives. AI doesn’t understand culture. It follows rules. Humans define them.

    Governance frameworks dictate how AI-generated outputs are reviewed, approved, and deployed. A payment processing module might require human sign-off for any AI-generated change. Low-risk UI updates could deploy automatically. These frameworks aren’t static. They evolve as trust in AI grows—or erodes. The group stressed governance as calibration. Too little oversight, and you inherit AI’s flaws. Too much, and you negate its benefits.

    Harness engineering didn’t emerge from theory. It emerged from necessity. The focus group’s participants weren’t futurists. They were practitioners who’ve seen AI-generated code fail in production. They didn’t ask if humans would remain in the loop. They asked how to keep them there effectively.


    Criticality: The Universal Criterion for Human Oversight

    Criticality = uncertainty + cost of change. That’s the formula the EuroPLoP 2026 group arrived at as the universal criterion for calibrating human oversight.

    Uncertainty isn’t just about unknowns. It’s about the kind of unknowns. Novel technology? Unclear requirements? Poorly understood domain? AI excels at well-defined problems. It struggles with ambiguity. The group cited systems with "subtle semantics" and "deeply entangled business logic" as examples where AI-generated solutions often miss the mark. These aren’t edge cases. They’re the norm in large-scale enterprise systems.

    Cost of change isn’t just financial. It’s risk. Regulatory fines? Reputational damage? Operational downtime? The higher the cost, the less room for error—and the less tolerance for AI’s probabilistic outputs. Legacy systems, with their technical debt and regulatory constraints, often fall into this category. AI can suggest a refactor. It can’t weigh the business impact of a failed migration. That’s a human judgement call.

    Criticality acts as a sliding scale. Low-criticality tasks—boilerplate code, repetitive patterns, isolated modules—can be safely automated. High-criticality decisions demand human oversight. The group didn’t just propose this as a guideline. They treated it as a structural necessity. AI can’t assess criticality. It doesn’t understand risk. It follows patterns. Humans must define where the line is drawn—and move it as needed.

    Brian Guthrie’s data shows AI-generated PRs are 18% larger. That’s fine for a low-criticality component. Unacceptable for a core payment processing module. The focus group’s participants didn’t just agree on criticality as a criterion. They described it as the missing framework in most AI-assisted development workflows today.


    Accountability: Why the Buck Still Stops with Humans

    Responsibility doesn’t emerge within AI systems. It’s assigned to humans. That’s not philosophy. It’s legal and ethical reality. The EuroPLoP 2026 group was unequivocal: accountability remains fundamentally human.

    AI can’t be sued. It can’t testify in court. It can’t explain its decisions to a regulator. When a security breach occurs—like the 2.74× higher XSS vulnerabilities in AI-generated code—someone must answer. That someone is the architect, the CTO, or the developer who signed off on the change. The group noted this reality shapes how organisations adopt AI. It’s not just about trust in the technology. It’s about trust in the people who govern it.

    The scenarios where AI fails aren’t just technical. They’re contextual. A security vulnerability might be a technical flaw, but its impact is reputational. A contract violation might be a compliance failure, but its consequence is legal. AI doesn’t weigh these trade-offs. It doesn’t understand consequences beyond technical correctness. The iSAQB’s perspective aligns: responsibility requires contextual judgement, not just technical execution.

    The group identified three key scenarios where accountability is non-negotiable:

    1. Security breaches: AI-generated code carries 75% more logic errors and 2.74× more XSS vulnerabilities. When a breach occurs, the question isn’t whether the AI made a mistake. It’s whether the human review process was adequate.
    2. Contract violations: APIs, data sharing agreements, and regulatory requirements often have subtle semantic constraints. AI doesn’t understand these. It follows patterns. Humans must ensure compliance.
    3. Long-term maintainability: Product debt—the gap between what the system does and what users need—still depends on humans watching, thinking, and caring about the commercial problem. AI doesn’t care. It optimises for what it’s told to optimise for.

    The group didn’t just identify these scenarios. They described accountability as the binding constraint on AI adoption. Organisations aren’t just asking what AI can do. They’re asking who will answer when it goes wrong. The answer isn’t changing. It’s still humans.


    The Structural Limits of AI in Architectural Decision-Making

    AI excels at repetitive, well-defined tasks. It struggles with ambiguity, long-term thinking, and ethical trade-offs. That’s not a limitation of current AI. It’s a structural constraint of the approach.

    The EuroPLoP 2026 group was clear: AI is a "power tool," not a replacement for human architects. It can generate code, suggest patterns, and optimise performance. But it can’t decide what to build, why to build it, or how to balance competing priorities.

    Ambiguity is the first hurdle. AI works best with clear requirements. But software architecture isn’t about clarity. It’s about navigating uncertainty. The group cited systems with "unclear boundaries" and "deeply entangled business logic" as examples where AI-generated solutions often fail. These aren’t edge cases. They’re the norm in large-scale enterprise systems. AI can suggest a microservice design. It can’t decide whether that design aligns with a company’s tolerance for operational complexity or its long-term roadmap.

    Long-term thinking is the second hurdle. AI optimises for immediate outcomes. It doesn’t consider future flexibility, technical debt, or evolving business needs. The group noted architectural decisions often involve trade-offs between short-term delivery and long-term maintainability. AI doesn’t weigh these trade-offs. It follows patterns. Humans must define the balance.

    Ethical trade-offs are the third hurdle. Privacy, fairness, unintended consequences—these aren’t technical problems. They’re human ones. AI doesn’t understand ethics. It doesn’t care about consequences beyond technical correctness. The group cited examples where AI-generated solutions introduced biases or violated privacy norms. These aren’t bugs. They’re failures of judgement. And judgement is a human domain.

    The data supports this. AI-generated PRs are 18% larger, suggesting less efficient or more fragmented solutions. That’s not just a technical flaw. It’s a symptom of AI’s inability to see the bigger picture. The group’s participants didn’t just observe this. They described it as the fundamental limit of AI in architecture. AI can execute. But it can’t decide.


    Guardrails: Why Humans Must Define the Rules

    Architectural guardrails—security policies, performance thresholds, compliance requirements—are intentional trade-offs. AI can enforce them. It can’t define them. That distinction matters.

    The EuroPLoP 2026 group was clear: the authoring of guardrails remains a human task. Guardrails require judgement. What’s the right balance between security and usability? How much technical debt is acceptable? What risks are worth taking? These aren’t technical questions. They’re business ones.

    The group cited a company’s tolerance for technical debt as an example. AI doesn’t understand debt. It doesn’t care about the cost of future changes. It optimises for what it’s told to optimise for. Humans must define what’s acceptable—and what’s not.

    AI also lacks the ability to anticipate edge cases. The group noted guardrails often emerge from past failures—security breaches, performance bottlenecks, compliance violations. AI doesn’t learn from history. It follows patterns. Humans must distil those lessons into rules.

    The group described guardrails as the interface between human judgement and AI execution. Humans define the rules. AI follows them. But the rules themselves—what they are, how they’re enforced, when they’re updated—remain a human responsibility.

    This isn’t just theoretical. The group’s participants described scenarios where AI-generated code violated guardrails—security policies, performance thresholds, compliance requirements. The issue wasn’t that the AI broke the rules. It was that the rules weren’t adequately defined or enforced. Humans must own that.


    Trust and Validation: The Human Feedback Loop

    Trust in AI-assisted architecture isn’t given. It’s earned. Incrementally. The EuroPLoP 2026 group described trust as a feedback loop—one requiring rigorous validation, controlled rollouts, and human oversight.

    AI-generated code isn’t deterministic. It’s probabilistic. Outputs vary, even with the same inputs. This variability demands human review. Automated testing isn’t enough. Someone must assess whether the output aligns with the intent.

    The group described a tiered approach to validation:

    • Low-criticality components: Automated testing, static analysis, compliance checks.
    • Medium-criticality components: Human review of AI-generated outputs, paired with automated validation.
    • High-criticality components: Full human oversight, with AI used as a "co-pilot" for exploration, not execution.

    This isn’t just about risk. It’s about calibration. The group noted organisations often start with strict oversight, then relax it as trust in AI grows. But that trust must be earned. It’s not a one-time decision. It’s an ongoing process.

    Feedback plays a crucial role. When AI-generated code fails, that failure must be analysed, understood, and fed back into the system. This isn’t just about fixing bugs. It’s about improving the AI’s understanding of the domain. Humans must own that loop.

    The iSAQB’s perspective aligns: technical excellence is necessary, but not sufficient. Human competence makes the difference—especially in complex and unpredictable situations. The group didn’t just describe trust as a goal. They described it as a process—one requiring human judgement at every step.


    Education: The Bottleneck in AI-Assisted Architecture

    The role of the software architect isn’t disappearing. It’s evolving. And education isn’t keeping pace.

    The EuroPLoP 2026 group was clear: universities and bootcamps aren’t teaching the skills that matter in AI-assisted architecture. Rote coding? Less important. Judgement? More important. Harness engineering? Almost non-existent in curricula.

    The group identified three key shifts in architectural education:

    1. From execution to judgement: Architects must evaluate trade-offs, assess risk, and make decisions under uncertainty. These aren’t technical skills. They’re human ones.
    2. From patterns to governance: Knowing design patterns isn’t enough. Architects must design the systems that govern AI’s use of those patterns.
    3. From code to context: Understanding business needs, regulatory constraints, and organisational culture is now as important as understanding technical constraints.

    These skills aren’t just missing from education. They’re often missing from industry. Many organisations still treat architecture as a technical discipline, not a human one. That’s changing—but not fast enough.

    The group described harness engineering as the emerging field for architects. It’s not just about designing systems. It’s about designing the systems that govern AI-assisted system creation. That’s a new skill set—one combining technical expertise with governance, risk assessment, and organisational design.

    The opportunity is clear. Architects who master AI governance will be in high demand. But the gap is real. The focus group’s participants described education as the bottleneck. Universities aren’t teaching this. Bootcamps aren’t covering it. Many architects are learning it on the job—through trial and error.


    The Path Forward: Automation’s Limits and Human Judgement’s Role

    AI will handle more of the "how. " That’s the central tension in AI-assisted software architecture—and the focus group’s core finding.

    The EuroPLoP 2026 participants didn’t just observe this tension. They proposed a framework for navigating it:

    1. Use criticality as a decision criterion: High uncertainty? High cost of change? Humans must decide. Low criticality? Automate.
    2. Invest in harness engineering: Build systems to govern AI-assisted development. Validation mechanisms, knowledge layers, governance frameworks—these aren’t optional. They’re structural necessities.
    3. Treat AI as an amplifier, not a replacement: AI can suggest, optimise, and execute. But it can’t decide, judge, or answer. Humans must define the rules—and enforce them.

    The group didn’t just propose this as a best practice. They described it as the only viable path forward. AI’s flaws—logic errors, security vulnerabilities, fragmented solutions—aren’t going away. They’re inherent to the approach. But they’re manageable—if humans remain in the loop.

    The most successful architectures won’t be those where AI does the most. They’ll be those where humans govern the best. That’s not a prediction. It’s a finding. The evidence is clear. The question isn’t whether humans will remain in software architecture. It’s how well we’ll adapt to their new role—and whether education will catch up in time.