IA 360
General Artificial Intelligence (AGI)

Creativity With AI: Measuring Novelty, Value, Process, and Authorship

Generating a work proves neither AGI nor creativity. A method for evaluating the artifact, process, selection, and provenance.

Admin IA360 4 min read AI-generated Leer en español
Creativity With AI: Measuring Novelty, Value, Process, and Authorship

Re-edited on July 30, 2026, this article removes a foundational confusion: an image generator, a language model, and AlphaFold are not examples of artificial general intelligence. They can produce unexpected or valuable results inside specialized systems. Calling them AGI does not help decide whether creativity occurred; it merely replaces a difficult question with a larger label.

“Can a machine be creative?” contains at least four questions: whether an artifact is novel, whether someone finds it valuable, which process the system followed, and to whom intent and responsibility are attributed. Those questions do not share one test. A work can be evaluated without claiming that a program had subjective experience, while an algorithmic contribution can be recognized without erasing the team that framed, selected, and presented the result.

Separate artifact, process, authorship, and impact

Novelty is always relative to a reference: new to the user, to the training set, to a genre, or to the documented history of a field? Value likewise depends on a criterion and community: satisfying a constraint, communicating an idea, opening an aesthetic direction, or serving a purpose. An output may be rare without value, or useful while repeating a known solution.

Process asks how the space of possibilities was explored: sampling, search, rules, evaluation, iterations, and selection. Authorship asks who set the problem, supplied data, designed the system, chose among variants, edited, and decided to publish. Impact adds provenance, labor, consent, concentration of power, and effects on the field. One “creativity” score erases those layers.

Human likeness should not become the definition either. A judge’s inability to distinguish origin may measure imitation under a particular test, not novelty or value. A perfect forgery would be hard to distinguish and would still have a problematic relationship with the copied work. Evaluation should examine both what was produced and the process that made it possible.

What generative models actually optimize

The 2014 paper on generative adversarial networks defined a game between a generator and discriminator: one tries to produce samples the other cannot distinguish from the data, while the other improves that distinction. The discriminator does not judge artistic value; it learns a statistical task. An image that appears to belong to a distribution is not thereby demonstrated to be original.

GPT‑3 was trained to predict tokens and evaluated across many tasks using examples and instructions in context. It can complete poems, stories, or code because those activities can be represented as sequences. Its objective contains no definition of aesthetic merit or intent. Prompt choice, temperature, tools, candidate count, and editing all belong to the observable creative system.

Variation does not guarantee meaningful diversity either. Thousands of samples may change surface details while remaining inside one template. Studying the process requires fixing the seed and version, retaining all variants rather than only the winner, and measuring distance and coverage through more than one representation. Later human selection may be the contribution that turns a distribution of drafts into a work.

How to measure novelty and value without pretending objectivity

Novelty begins with proximity searches against a declared corpus: exact matches, fragments, structures, and semantic neighbors. This does not prove absence of all influence, but it can detect repetition. Extracting Training Data from Diffusion Models showed that diffusion models could memorize and emit training images under certain conditions. “The model generated it” is therefore not sufficient evidence of novelty.

Value is assessed with a rubric tied to the brief: satisfaction of constraints, coherence, relevant surprise, usefulness, elaboration, or response from a defined audience. Several judges review blindly where possible, and their degree of agreement is reported. Disagreement is not always noise to remove; it can reveal that a field does not share one criterion or that the work opens a new interpretation.

A comparison needs baselines. The system is tested against retrieval of existing works, templates, random variants, a previous method, and human work under stated time and resources. Both best cases and the full distribution are measured. Showing one piece selected from ten thousand tests the system plus a vast selection budget, not the typical quality of one generation.

Two phases should also be separated. In a divergent phase, the questions are how many substantially different directions appear and what it costs to find them. In a convergent phase, evaluation asks whether the process recognizes constraints, compares alternatives, and improves a proposal after feedback. One system may generate variety while being unable to choose; another may optimize a metric while removing surprise. Stating which phase a machine performed and which a person performed prevents the model from receiving credit for the entire cycle.

Three cases with three different boundaries

The Next Rembrandt was an ING project involving Microsoft, TU Delft, Mauritshuis, and Rembrandthuis. A multidisciplinary team analyzed works, defined features for a portrait, and produced a 3D-printed image. It was not an AGI that decided to become a painter: it was a human and technical installation designed to imitate documented properties of one artist. Its interest lies in the procedure and collaboration, not in relabeling it as an autonomous mind.

AlphaFold predicts protein structures from sequences and related data. The paper does not claim that the system “conceived new proteins,” as the original article stated. Structure prediction and designing a sequence with new properties are different tasks. A scientifically useful output can expand research without the model framing the question, interpreting the mechanism, or choosing the next experiment.

FunSearch combined a language model, an automatic evaluator, and an evolutionary process to search for programs solving mathematical and algorithmic problems. The loop is essential: the model proposes code, the evaluator assigns a verifiable score, and the best proposals feed later iterations. The result demonstrates useful novelty inside defined problems; it does not establish general creativity outside that mechanism.

An auditable record of assisted creation

The record starts with the brief and success criterion. It lists authors and roles; sources and permissions; model, version, and date; prompts, tools, seeds, and parameters; total candidates; rejection rules; editing; evaluators; and results. Representative failures are retained. This trace allows contributions to be attributed and the process repeated without reducing it to the model’s name.

Evaluation should be repeated outside the group that built the system. A person who designed the metric may inadvertently have made it an easy target to exploit, while someone who chose the examples knows the answers. Independent review searches for omitted similarities, tests other audiences, and checks whether the result retains value once the process is disclosed. Judgment may change with that context, and the change is part of the finding.

Provenance metadata helps but does not settle judgment. The C2PA specification defines signed manifests for recording provenance and actions on digital content. It can document that a tool created or modified a file when the workflow incorporates it. It does not determine whether a work is original, true, valuable, or who owns rights; those questions require evidence and, where appropriate, jurisdiction-specific legal review.

Finally, state a narrow conclusion: “this system produced variants that defined judges valued under this rubric and these references,” not “the machine is creative.” The transferable skill is separating novelty, value, process, and authorship and requiring evidence for each. That makes it possible to appreciate a new collaboration without inventing AGI, consciousness, or autonomy where only an artifact has been observed.

This article was produced with artificial intelligence under human editorial oversight.

Share this article

This website uses cookies to improve the browsing experience. Cookie policy.

↑↓ navigate ↵ open esc close