Skip to main navigation Skip to search Skip to main content

Codified collaboration: reinforcement learning with verifiable feedback as a mechanism for human–AI co-creation in generative intelligent design

Research output: Journal contributionsJournal articlesResearchpeer-review

Abstract

Generative intelligent design systems typically remain static: they generate plausible artifacts but do not adapt through use or internalise expert design knowledge in a principled way. To address this limitation, we propose reinforcement learning from verifiable feedback as a framework for codified human–AI collaboration, in which domain expertise is formalised as algorithmic verifiers that evaluate generated designs against explicit structural and logical rules and convert rule compliance into reinforcement signals. We instantiate RLVF on the engineering task of functional decomposition, where designs are represented as typed function–flow graphs governed by verifiable structural constraints. Using group relative policy optimisation, a Llama-3.1-8B model is trained exclusively from verifier feedback without labelled output data. Across reinforcement learning cycles, correct output formatting improves from 18% to 100%, fully connected functional structures improve from 10% to 100%, and error-free decompositions improve from 4% to 100%, exceeding a supervised fine-tuned baseline trained on 257 labelled examples. Human evaluation shows that RLVF-trained models achieve higher perceived logical coherence than supervised models, while exhibiting reduced structural diversity. These results demonstrate that verifiable feedback can replace human annotation in rule-governed design tasks and enable correctness-grounded generative design systems based on codified expert knowledge.

Original languageEnglish
JournalJournal of Engineering Design
Number of pages26
ISSN0954-4828
DOIs
Publication statusAccepted/In press - 2026

Bibliographical note

Publisher Copyright:
© 2026 The Author(s). Published by Informa UK Limited, trading as Taylor & Francis Group.

Research areas and keywords

  • codified expertise
  • functional decomposition
  • generative intelligent design
  • Human–AI collaboration
  • reinforcement learning from verifiable feedback (RLVF)
  • Engineering

ASJC Scopus Subject Areas

  • General Engineering
  • Engineering(all)

Fingerprint

Dive into the research topics of 'Codified collaboration: reinforcement learning with verifiable feedback as a mechanism for human–AI co-creation in generative intelligent design'. Together they form a unique fingerprint.

Cite this