Why I Am Not Building Another Image Generator

Published On: June 28th, 2026Last Updated: September 14th, 20262563 words12.8 min readDaily Views: 1Total Views: 11

There are already enough buttons that say Generate.

The largest technology companies have spent extraordinary amounts of money training image models, building infrastructure around them, moderating their use, and integrating them into increasingly polished platforms. An independent builder cannot realistically reproduce that work at the same scale—and wrapping those existing models in another interface would not create a new image generator anyway.

It would create another layer.

Another account. Another subscription. Another set of usage limits. Another product dependent on external pricing, policies, model changes, regional availability, and API access.

That is not the gap I am interested in filling.

I am building a Visual Register by Mithaq Praxis because the problem often begins before anyone presses Generate.

It begins when a person knows they want an image, but has not yet decided what that image is supposed to do.


The problem is not always the model

A recent example began with something very ordinary: a series of playful summer scenes featuring four adult women.

The intended images were not complicated in concept. Friends running toward a pool. A chaotic volleyball game. A floating breakfast. A night swim beneath terrace lanterns.

The generator repeatedly rejected or misunderstood parts of the request.

Some of this came from the platform’s safety interpretation. Some came from the combination of group scenes, swimwear, action, and unspecified ages. Some came from the fact that several scenes had been submitted together. Another generator, thread, or platform might have interpreted the exact same language differently.

That is one of the realities of generative-image work: there is no universal sentence that behaves identically everywhere.

Different models have different training, safeguards, context handling, reference-image systems, visual strengths, and tolerance for ambiguity. A prompt that works well in one platform may fail in another. Even within the same platform, a long-running conversation containing established character references may behave differently from a new blank session.

The answer is not simply to write a much longer prompt.

Length is not the same as clarity.

A prompt can contain hundreds of words and still fail to answer the most important questions:

  • What is this image for?
  • What is actually happening inside the frame?
  • Which details are essential?
  • Which details may vary?
  • What should remain consistent with the creator’s wider body of work?
  • What would make the result unsuitable, even if it looked attractive?

Those questions belong to visual direction, not merely prompt construction.


I do not want to build another wrapper

It would be technically possible to connect several image-generation APIs to a new interface, charge users for credits, and market the result as a new creative platform.

But I would still be relying on someone else’s model to produce the image.

The cost of generating, storing, moderating, and delivering those images would eventually be passed to the user—often someone who already pays for access to the original generator.

The product would also inherit every external dependency:

  • pricing changes;
  • rate limits;
  • model retirements;
  • policy changes;
  • regional restrictions;
  • capability differences;
  • moderation failures;
  • API deprecations.

More importantly, building around the Generate button would place the emphasis in the wrong place.
It would suggest that the most valuable part of the process is the moment the model renders the pixels.

I do not believe that it is.

The model may render the image. The human still has to decide what deserves to be rendered.


The missing layer is visual pre-production

Before a film scene is shot, people make decisions.

They decide who is present, where each person stands, what the scene communicates, where the light comes from, what the environment says about the story, and what the viewer should notice first.

  • Before a designer produces a campaign, someone has to establish the audience, message, visual hierarchy, brand language, required assets, and practical format.
  • Before an illustrator begins a recurring world, characters and environments develop visual rules. Colours recur. Clothing belongs to a particular culture, season, or role. Certain objects carry meaning. Some details must never drift.

Generative-image platforms did not make those decisions irrelevant.
They made it easier to hide them.

A person can now type several aesthetic words, produce twenty images, select the most attractive one, and only afterwards attempt to explain why it belongs to the project.

That is not always deliberate visual work. Sometimes it is simply visual luck followed by retrospective justification.

The Visual Register begins earlier.

It asks the person to decide what they are making before the system helps organise how it may be rendered.


The human must write the first line

Every new visual brief in The Visual Register begins with a human-written intention.

  • It does not need to be elegant.
  • It does not need to use professional cinematography language.

It may be as simple as:

Four friends run toward a pool together, caught just before they jump.

That one line establishes something important: the image begins with a human decision.

The system can then help the user clarify the scene.

  • Are all four people adults?
  • Does each person have a distinct action?
  • Is the mood competitive, affectionate, chaotic, or calm?
  • Is the camera close to the water or looking down from the terrace?
  • Is the image intended for a blog thumbnail, a full-page editorial, or a vertical social post?
  • Should the clothes, architecture, colours, or character features match an existing visual world?

The tool may provide structure and craft vocabulary around the person’s idea. It should not quietly invent the underlying intention and then ask the user to approve the machine’s guess.

That distinction matters.

There is a difference between human-led work and work that is merely human-fronted.

  • In human-led work, the person decides what the image means.
  • In human-fronted work, the machine proposes the meaning and the person accepts or rejects it.

The Visual Register is being designed to keep the authorship line on the human side.


The Visual Bible is the centre of the product

The main object inside The Visual Register is the Visual Bible.

A Visual Bible is a persistent source of truth for a project, world, brand, publication, company, or campaign.

It holds the decisions that should remain stable across multiple images:

  • visual identity;
  • colour palette;
  • recurring characters;
  • attire;
  • architecture;
  • environments;
  • materials;
  • motifs;
  • personal easter eggs;
  • logos;
  • product details;
  • cultural boundaries;
  • continuity locks;
  • avoid lists;
  • reference material.

I already work this way informally.

I have predefined visual systems for three different worlds: Bayt al-ʿAhd, Algorithm Atelier, and Mithaq Praxis.

Each has its own architecture, emotional register, palette, recurring objects, attire, spaces, and visual atmosphere.

  • The Bayt remains recognisable as the Bayt whether the image is photorealistic, painterly, illustrated, sepia-toned, doll-like, or rendered as a graphic novel panel.
  • Mithaq Praxis remains recognisable as Mithaq Praxis whether I am showing a research library, a systems laboratory, a corporate website image, or a quiet editorial room.

The rendering style can change.

The underlying visual identity remains coherent because the deeper decisions already exist.
That is what the Visual Bible preserves.

It does not force every image to look identical. It stores the identity beneath the treatment.


Store what should remain. Ask what is new.

This became one of the central product laws:

Store the decisions that should remain. Ask only for the decisions that are new.

A user should not have to rewrite the same character description, palette, architecture, logo rules, or wardrobe restrictions for every image.

Those belong in the Visual Bible.

But the Bible should not decide what happens in the present scene.

The user still has to provide what is new:

  • the current action;
  • the purpose of the image;
  • the emotional tone;
  • who is present;
  • what each person is doing;
  • where the viewer’s attention should go;
  • what the image must communicate.

This is how the tool can reduce unnecessary repetition without automating away meaningful thought.

The Visual Bible is also optional.

Not every image belongs to a recurring world. Someone may need one visual for one presentation, announcement, article, or personal project.
They should still be able to use the brief composer without first constructing an entire identity system.

The Visual Register therefore has two layers:

  • A reusable Visual Bible for continuity when continuity is needed.
  • A visual brief composer for every new image, whether a Bible is attached or not.

The corporate problem is not “bad prompting”

The same gap exists inside organisations.

  • A company may announce that employees are permitted to use generative AI, while providing no shared method for doing so.
  • One employee may use a three-word prompt.
  • Another may provide a detailed visual specification.

A designer may understand framing, hierarchy, lighting, audience, negative space, and brand continuity. A manager may believe that everyone is performing the same task because they are all using the same generator.

The result is predictable:

  • inconsistent quality;
  • brand drift;
  • unsuitable images;
  • confidential information entering prompts;
  • unclear responsibility;
  • unrealistic expectations of designers;
  • departments producing visuals that do not belong to the same organisation;
  • polished images that fail the communication goal.

The problem is not simply that some employees are “bad at prompting.”
The organisation has not created a shared visual language.

It may have a traditional brand guide, but that guide was not necessarily designed to answer the practical questions that appear when staff generate new scenes:

  • What does the company’s research laboratory look like?
  • How are people represented?
  • What visual clichés should be avoided?
  • Which products must be depicted accurately?
  • How much space must remain for text?
  • Which logos may be used?
  • What tone is appropriate for recruitment compared with a technical report?
  • Who reviews the final image?

The Visual Register cannot solve every organisational problem around AI.
It can provide a common starting point.

  • A company can create and govern its own Visual Bible. Staff can use the same brief composer while inheriting approved brand rules, assets, and restrictions.
  • The person requesting the image still has to state what it is meant to communicate.

The Bible provides continuity. It does not replace responsibility.


Designers do not merely press a button

One of the most damaging assumptions around AI-assisted design is that faster generation means the design process has disappeared.

A model can produce a polished-looking image in seconds.
That does not mean the image is correct.

A designer may still need to:

  • interpret an unclear request;
  • identify the audience;
  • select an appropriate concept;
  • maintain brand consistency;
  • create several directions;
  • reject visually attractive but unsuitable outputs;
  • correct anatomy or product details;
  • rebuild typography;
  • protect space for layout;
  • check accessibility;
  • review cultural implications;
  • adapt the image to multiple formats;
  • combine generated and manually created elements;
  • retouch the final asset;
  • explain why one version works and another does not.

The speed of rendering does not remove the need for judgment.
It may change where that judgment occurs.

A useful visual system should make that work more legible, especially to non-designers.

A lightweight decision record might show:

  • Draft one rejected because the emotional tone was too playful for the audience.
  • Draft two rejected because the product proportions were inaccurate.
  • Draft three revised because there was insufficient negative space for the headline.
  • Final version selected after brand and accessibility review.

That record does not exist to inflate the process.

It exists to show that a finished concept is a body of decisions, not merely the output of a button.


Reading and vocabulary are part of the work

Visual direction depends on the ability to notice and name distinctions.

A person may know that an image feels wrong without knowing whether the problem is:

  • harsh overhead light;
  • weak hierarchy;
  • generic posing;
  • an unsuitable lens feeling;
  • excessive symmetry;
  • a mismatch between mood and colour;
  • a crowded background;
  • the wrong type of quiet.

Language helps us examine those differences.

That is why The Visual Register will include a learning layer: short explanations, examples, glossaries, and links to useful external resources such as dictionaries, thesauruses, and visual-reference material.

  • The purpose is not to make the system write the entire scene for the user.
  • The purpose is to help the user build a vocabulary they can carry beyond the tool.

A button labelled Improve my prompt can easily replace the user’s intention with a machine-authored version containing choices they never made.

A more human-led system asks:

You wrote “dramatic lighting.” Which kind do you mean?

Then it explains several possibilities and lets the person decide.

The tool should help people become more capable, not more dependent.
It should require deliberateness quietly, through structure.

The essays can carry the argument. The interface does not need to lecture.


This is not an authorship certificate

The Visual Register can preserve a body of decisions:

  • the original scene seed;
  • the attached Visual Bible;
  • references;
  • selected visual choices;
  • rejected versions;
  • revision notes;
  • final selection;
  • later human edits.

That record may be useful for provenance, internal review, creative documentation, and explaining the work behind a finished image.
It should not claim that using the tool automatically guarantees copyright or legal ownership.

Those questions depend on jurisdiction, process, source material, and the nature of the human contribution.

The honest promise is simpler:

Preserve the human decisions behind the work.

A person who has directed the intention, constructed the visual world, selected the references, rejected unsuitable results, corrected the work, and shaped the final asset has a visible creative process.

The tool documents that process.
It does not adjudicate it.


What I am actually building

The first version of The Visual Register is deliberately smaller than the full system it may eventually become.

Its initial spine is:

A Visual Bible builder

For reusable visual identity, continuity rules, environments, characters, attire, motifs, references, and restrictions.

A visual brief composer

Beginning with a human-written scene intention and expanding only as much as the user needs.

A clean human-readable brief

Something that can be copied into a generator, handed to a designer, attached to a project request, or stored alongside the resulting work.

Lightweight decision notes

Enough to record why a version was revised, rejected, selected, or approved.

It will not begin as an image generator.
It will not maintain a brittle library of hand-written adapters for every platform.
It will not promise identical results across models.
It will not replace designers.
It will not automatically invent the user’s scene and call that human-led creation.

Those limits are part of the design.


Before the image

The Visual Register began from a practical frustration, but the deeper question is cultural.

  • What habits do we want generative tools to normalise?
  • Do we want people to become increasingly passive—typing less, reading less, deciding less, and approving whatever arrives?
  • Or can we build tools that use automation without treating thought as an inconvenience?

I am interested in the second path.
I do not believe every image needs a thesis, a production team, or a twenty-field form.

A person may write one sentence and make only a few choices.
But those choices should still belong to the person.

The aim is not maximum complexity.
It is visible intention.

  • The generator may produce the pixels.
  • The platform may provide the model.
  • The Visual Register will exist for the work that happens before that moment: the world, the purpose, the language, the continuity, the boundaries, and the decisions that make an image belong to someone’s wider practice.

Because good image-making does not begin when the button is pressed.
It begins when a person decides what matters inside the frame.

© 2026 • MITHAQ PRAXIS • CC BY-NC-ND 4.0 Unless Otherwise Stated.