Structured data is a small block of code that tells a machine what a page is about, in a form it does not have to guess at. A person reads your page and understands it. A crawler reads the same page as a wall of text, then works out for itself which words are the phone number, which are the author, and which merely mention a rival.
That guessing is where things go wrong. AI answers pull facts from many pages at once, so anything ambiguous either drops out or comes back wrong. Markup removes the ambiguity. It will not write your content, and it will not lift a weak page above a strong one, yet it makes a good page much easier to quote.
What structured data actually is
Two things hide behind the phrase. First, a shared vocabulary: a public list of types and properties such as organisation, article, product and opening hours. Second, a syntax, which is simply how you write that vocabulary into a page.
The vocabulary belongs to nobody in particular. It exists so that different engines can read the same labels, and that shared reading is the value. You are not inventing labels for your own site, you are borrowing labels the machines already know.
None of it shows to a reader. Your visitor sees a headline and an address, while the markup sits quietly in the source, stating that the headline is the article name and the address is a postal address.
Why AI answers lean on structured data
A language model assembles an answer out of fragments. It wants facts it can lift without much interpretation, and it favours facts that agree with each other across sources.
Markup helps at exactly that step. A labelled fact needs no interpretation at all. When the visible page and the block underneath it say the same thing, a machine gets a second, cleaner copy of the claim it was about to make.
Why an answer picks any source in the first place is a wider question, and our guide to generative engine optimization works through that mechanism. This article stays on the markup layer, because that layer is concrete and you can actually finish it.
What a machine does with your labels
Three jobs, roughly. Extraction, where it lifts the fact off the page. Matching, where it decides your organisation is the same one it met elsewhere. Display, where a search engine turns the fact into a richer result.
The middle job matters most here. Two businesses share a name more often than owners expect, so anything that pins your identity down lowers the chance of a model crediting your work to somebody else.
Pick JSON-LD and ignore the other formats
Three syntaxes exist in practice. Microdata and RDFa weave attributes through your HTML, so the labels and the layout end up tangled together. JSON-LD sits apart, in one block of its own.
Use JSON-LD unless your platform forces your hand. It survives a redesign, and one person can edit it without touching a template. Tangled markup breaks quietly when somebody moves a heading, and nobody notices for months. Search engine documentation is the place to confirm the current preference, because that guidance is theirs to change.
The structured data types worth your time
The vocabulary runs to hundreds of types. Most businesses need a handful, and piling on more does not make the signal stronger.
| Type | What it settles |
|---|---|
| Organization | Who you are, plus the profiles that genuinely belong to you. |
| LocalBusiness | Address, hours and telephone number when customers visit in person. |
| Article | Headline, author and dates, so a post is more than a page. |
| Product | Name, availability and real reviews if you sell online. |
| FAQPage | Question and answer pairs, once they appear on the page itself. |
| BreadcrumbList | Where a page sits, especially on a deep site. |
Begin with organisation and article, since nearly every site has both. Then add product or local business only if they genuinely describe you. Everything after that is optional.
Structured data has to match the visible page
This rule outranks the rest. Whatever the markup claims, a reader must be able to find the same fact on the page itself.
Marked-up hours that differ from the hours on the page cause real damage. An author name nobody can find anywhere does the same. Review stars for reviews that never happened are the clearest breach, and search engines publish rules against exactly that. Read the current guidelines yourself rather than trusting a summary in an article.
Assume the two get compared. When the markup and the page disagree, a machine has no reason to trust either. You never see the moment that happens, so the fix is to keep them in step. Honest markup is the cheap version of trust.
Structured data and the entity behind the business
An organisation block does a job the rest of your content cannot. It states your name, your website and the profiles that belong to you, in one place, in a form a machine can match against other sources.
That matching sits at the heart of entity SEO, and markup is one of its steadier inputs. Name, website address, logo and the profile links that genuinely belong to you: that is most of the block. Where your name appears in three slightly different forms across the web, a clean organisation block hands every crawler one version to trust.
Does structured data reach the assistants?
Partly, and honesty serves you better than confidence here. Some assistants fetch pages live, others lean on a search index that has already parsed your markup, and the mix changes as the products change.
Nobody outside those companies can tell you how much weight each system puts on it. What we can say is narrower and still useful. Clean labels make your facts unambiguous, and unambiguous facts travel further than implied ones.
So treat markup as insurance rather than as a lever. If the reading improves, your labels are already waiting. Worth remembering too: the words on the page still do most of the work, which is the subject of writing content for AI search.
The FAQ markup trap
Question markup looks like free space in the results, so people bolt it onto pages that were never FAQs.
Two problems follow. First, the way search engines display question markup has shifted over the years, so any visible rich result is a bonus rather than a plan. Second, the questions must appear on the page in full, where a reader can see them, or the markup contradicts the page and the previous rule applies.
Used honestly, question and answer pairs still earn their place. They map neatly onto the way people ask an assistant things, which is why they come up so often in answer engine optimization.
What structured data will not do
Markup is no substitute for relevance. Nothing in the vocabulary makes a page a better answer, so a thin page with immaculate markup stays thin. If you are hoping labels rescue weak content, spend the time on the content instead.
Nor will it force an AI system to cite you. Citations come from being quotable and useful, which is the ground earning AI citations covers. Markup only makes the quoting easy once the content deserves it.
And it cannot fill a gap in your answers. Where a question has no answer anywhere on your site, no label invents one for you.
Testing structured data, then testing it again
Testing takes a minute, and the common validators are straightforward to use. Paste in the address of a page, read what the tool recognised, then fix whatever it could not.
Testing once is the mistake. A theme update, a plugin change or a rebuild can drop the block completely, and nothing on the visible page changes when it happens. So put a repeat check in the calendar, quarterly at least, and test templates rather than every page.
What to test, in order
- Homepage first, because the organisation block usually lives there.
- One blog post, since a single template covers every other post.
- A service or product page, if you sell through the site.
- Anything rebuilt recently, as rebuilds are where markup disappears.
How much structured data a small site needs
Less than most owners assume. Markup lives in templates, so you write it once per template instead of once per page.
A worked example, with round numbers picked to keep the arithmetic obvious rather than to describe a typical site. Say a site has 60 pages: 1 homepage, 5 service pages, 4 other pages and 50 blog posts. The posts share one template and the service pages share another, so 2 templates cover 55 of those pages. That leaves 5 blocks to write by hand, which is an afternoon of work rather than a project.
The real cost sits in the upkeep rather than the build. Every time your hours, your name or your author list changes, somebody has to change the markup too, and that is a habit rather than a task.
Where to start with structured data this week
- Check what your platform already outputs, because many add markup silently.
- Fix the organisation block first, since your identity holds everything else up.
- Add article markup to the blog template, then leave it alone.
- Validate three pages and write down what came back.
- Diary a repeat test for after the next redesign, when things break quietly.
None of this work is glamorous. Structured data is plumbing, and plumbing only draws attention when it fails. Still, it is one of the few jobs in AI search with a clear end, a small cost, and a result that does not depend on somebody else feeling generous.
Frequently asked questions
What is structured data, in plain words?
Structured data is a block of code on a page that labels the facts a machine would otherwise have to interpret. The labels come from a public vocabulary that search engines agree to read. A visitor sees no difference at all. A crawler sees your business name, your opening hours or your author clearly marked, rather than buried inside a paragraph it has to parse for itself.
Which markup format should I use?
JSON-LD, in almost every case. It sits in a single block instead of weaving attributes through your HTML, so a redesign is far less likely to break it. Microdata and RDFa still work, yet they tie your labels to your layout, which makes both harder to change later. If your platform already outputs one of the older formats correctly, leave it alone and spend the time elsewhere.
Does adding markup improve my rankings?
Not directly. Markup does not make a page more relevant to a question, so nothing climbs on the strength of it alone. What it can do is earn a richer result in some cases, make your facts easier for a machine to reuse, and settle any doubt about who you are. Those are real gains, but they are clarity gains rather than ranking ones.
Can I add this without a developer?
Often, yes. Many content platforms carry a plugin or a built in setting that writes the common types for you, and a generator can produce a block you paste into the page. The judgement is the hard part rather than the code. Choosing the right type, filling only the fields you can verify on the page, and testing the result all matter more than the syntax. A developer helps most on a custom build.
What happens if my markup and my page disagree?
You lose the benefit, and you risk more than that. Search engines publish rules that treat structured data contradicting the visible page as a quality problem, and the penalties they describe can reach past the single page. Check the current guidelines rather than working from a summary. The rule underneath is simple. Every fact inside the markup should appear somewhere a reader can see it, in the same form. Review stars for reviews that never happened are the clearest example.
How often should I test structured data?
Quarterly suits most sites, plus a test after any redesign, theme update or platform migration. Structured data disappears silently, since nothing on the visible page changes when a block vanishes. Test one page per template rather than every page, because templates are what actually break. Write down what the validator said each time, so next quarter you can see whether anything moved.
Is FAQ markup still worth adding?
Yes, when the page really is a set of questions and answers. Treat any rich result as a bonus, because the way search engines display question markup has changed more than once. The lasting value is different. Structured data on a question block hands an assistant clean pairs to quote from, which is closer to how people phrase things when they ask. Never mark up a question a reader cannot see.
Do language models read the markup directly?
Nobody outside those companies can say, and the honest answer is that it varies. The practical move is to stop depending on the question. Put every fact you care about into the visible text as well as into the structured data, so a system that only reads rendered words still finds it. Markup then becomes a second clean copy rather than the only copy. That holds however the reading works this year.
How do I know if structured data is doing anything?
Watch three things over a quarter. First, whether a validator still reports a clean block after site changes. Second, whether any richer results appear for the pages you marked up. Third, whether assistants describe your business correctly when you ask them about it. None of those gives you a clean attribution line, so treat this as upkeep you keep in order rather than a campaign with a monthly number.