makefield

Week 8 · the sixth dimension

Structured data: machine-true, and why schema isn't enough

Schema is the clearest signal you can hand an AI about what you are. It's also routinely oversold. Here's what it does — and what it can't.

The clearest signal

Telling a machine what you are, explicitly

Most of what an AI knows about you it has to infer from prose. Structured data — schema.org markup — lets you state it explicitly, in a form built for machines: a SoftwareApplication with its category, pricing, and rating; an Organization with its name, logo, and sameAs links to your authoritative profiles; FAQPage markup on your answers. It's the highest-confidence signal you can hand a system, because it removes the guesswork behind a vague or competing description.

Why schema isn't enough

A clarifier, not a magic lever

Here's where a lot of advice oversells. Schema is a clarifier — it helps a model parse what you are. It does not make the web's account of you correct, and it does not force a recommendation. If the most-cited description of you is wrong, schema won't override it; if you're absent from the sources an engine trusts, schema won't put you there. This is exactly the kind of single-fix overpromise that makes most AEO checklists wrong — schema is necessary hygiene, not a sufficient strategy. Do it, then keep going.

Machine-true

The markup has to match reality

The one rule that matters most: structured data must be true and current. Schema that contradicts your page, or that hasn't been updated since launch, is worse than none — some models distrust an entity whose explicit signals don't match its content. "Machine-true" means the markup says exactly what the page says, and both say what's actually so. Adding schema as decoration — markup for markup's sake, copied from a competitor, never maintained — is theatre, and a careful model can tell.

The newer surfaces

llms.txt and the machine-readable frontier

Beyond schema, a newer convention is appearing: files like llms.txt and /llm-info pages that hand AI systems a plain, curated summary of a site. It's promising and worth doing — but adoption is still uneven and the payoff unproven, so treat it as an addition after the basics, not a substitute for them. The order holds: clean, true content first; explicit schema second; the emerging surfaces third.

Next — Citation & authority: the sources an AI trusts to describe you, most of which aren't your own site. New here? Start with the overview.

Common questions

Questions this lesson answers

Does schema markup help AI understand my business?

Yes, in one specific way: it removes guesswork. Most of what a model knows about you it has to infer from prose, and schema.org markup lets you state the same facts explicitly, in a form built for machines. An Organization entry carries your name, logo and the sameAs links to the profiles you consider authoritative; a SoftwareApplication entry carries the category, pricing and rating; FAQPage markup carries your answers as answers. That is the highest-confidence signal you can hand a system about what you are, which matters most when the descriptions of you in circulation are vague or competing.

Is schema markup enough on its own to fix how AI describes my company?

No, and this is where a lot of advice oversells. Schema is a clarifier: it helps a model parse what you are. It does not make the web's account of you correct, and it does not force a recommendation. If the most-cited description of you is wrong, schema will not override it; if you are absent from the sources an engine trusts, schema will not put you there. It is necessary hygiene rather than a sufficient strategy, which is also why the single-fix checklist version of this advice fails. Add the markup, then keep going through the other dimensions.

Can wrong or outdated schema markup hurt me?

It can, and that makes stale markup worse than none at all. The rule is that structured data must be true and current: the markup says exactly what the page says, and both say what is actually so. Schema that contradicts the page it sits on, or that has not been touched since launch, gives a model two conflicting explicit signals, and some systems respond by trusting the entity less. Markup copied from a competitor and never maintained is the common version of this.

Do I need an llms.txt file?

It is worth doing, but after the basics rather than instead of them. Files like llms.txt, and llm-info pages, hand AI systems a plain, curated summary of a site, and the convention is spreading, but adoption is still uneven and the payoff is unproven. The order that holds: clean, true content first, explicit schema second, the emerging machine-readable surfaces third. A curated summary of a site whose underlying facts are vague inherits the vagueness.