Data, SEO & AI Search · Practical guide
Schema markup
What schema markup is, which types still earn a visible result, the two that quietly stopped in 2023, and why structured data matters more for AI answers than it now does for rich snippets.
The short answer
What is schema markup?
Schema markup is a block of labelled facts you add to a page so software does not have to guess what the page is about. It does not change a word of what a visitor reads. It states, in a shared vocabulary the search engines agreed on in 2011, that this page is an article, that a named organisation published it, that the business trades in Sydney and that these are its opening hours. Google recommends adding it as JSON-LD, a small block of structured text in the page source.
Key takeaways
Schema markup labels facts a page already contains. It is not new content, and adding it is not a ranking tactic.
Google supports a small set of types with a visible result. The schema.org vocabulary is far larger, and most of it earns you nothing on the results page.
Question markup and step-by-step markup, two of the most recommended types on the web, stopped producing results for normal businesses in 2023.
The value now sits in the unglamorous types: who you are, where you trade, who wrote this and where it sits on your site.
The mechanics
What schema markup actually is
A search engine reading an ordinary page is doing inference. It sees a phone number and infers a business; it sees a date and infers a publication; it sees a suburb and infers a service area. It is usually right, and when it is wrong there is nothing you can do about it, because you never told it anything. Markup is the part where you tell it.
The vocabulary itself is schema.org, published in 2011 by the major search engines so there would be one agreed way to say "this is an organisation" rather than four. The vocabulary is enormous and deliberately so, covering everything from a bus route to a medical trial. What matters commercially is much narrower: the handful of types Google will actually do something with.
There are three ways to write it. JSON-LD is a self-contained block that sits in the page source and touches nothing else; microdata and RDFa weave attributes through your visible HTML. Google recommends JSON-LD, and so do we, for the boring reason that it can be changed without anyone touching the layout.
| Format | Where it lives | The practical difference |
|---|---|---|
| JSON-LD | One block in the page source | Independent of the layout. A designer can rebuild the page and the markup survives untouched. Google's recommended format. |
| Microdata | Attributes on your visible HTML | Tied to the markup around it. Change the template and you can silently break it. |
| RDFa | Attributes on your visible HTML | Same fragility as microdata, with a heavier syntax. Rare on commercial sites now. |
Adding markup is not a ranking factor in the way people hope. Google has been consistent that structured data helps it understand a page and makes the page eligible for certain result types. Eligible is the operative word: it is a door, not a lift.
The bit nobody updated
What Google ignores, and what it stopped showing
Here is the part that makes most schema advice out of date. In August 2023 Google announced that frequently-asked-question rich results would be shown only for well-known, authoritative government and health websites. For every other site they would no longer be shown regularly. At the same time, step-by-step how-to rich results were limited to desktop. In September 2023 they were removed from desktop as well, and Google deleted the how-to documentation entirely.
What changed in 2023, and what still earns a visible result. Source: Google Search Central, "Changes to HowTo and FAQ rich results", August 2023, and Google's documentation update log, September 2023.
Neither has returned. Yet both are still recommended on page after page of SEO advice, sold as free real estate on the results page, and still shipped by plugins by default. A small business paying for question markup on twelve pages in 2026 is paying for a feature that was withdrawn three years ago.
That does not make the markup harmful, and it is not why we avoid it. We simply do not ship a type that cannot pay for the code it adds, and we hold a harder line on a second group: ratings, reviews and prices that a business writes about itself. Markup that asserts a five-star average nobody can verify is the fastest way to lose the trust that the rest of the markup is there to build. On our own site a page will not ship if it carries a rating, a review or a price we cannot stand behind, and that rule has caught more mistakes than it has cost us features.
What to actually add
The types worth an hour of your time
For a normal Australian business, four types carry almost all of the value, and a fifth if you sell things. None of them is exciting. All of them answer a question a machine would otherwise guess at.
Organisation
Your legal name, logo, contact details and the profiles that belong to you, stated identically on every page. This is the one that stops an engine confusing you with a similarly named business.
Local business
Address, opening hours and the areas you genuinely serve. For a Sydney business this is the difference between being understood as local and being understood as national.
Article
Author, publication date, last modified date and publisher. It is what lets an engine tell a maintained page from an abandoned one.
Breadcrumb
Where the page sits in your site. Cheap to add, and it shapes the path shown under your result.
If you sell products, product markup remains fully supported and genuinely worth the effort, because stock status and shipping detail change what a shopper sees before they click. Everything beyond these is a judgement call, and the honest test is the one Google publishes: does this type have a documented search feature attached to it, and does your page truly contain what the type claims?
The second half of that test is where most sites fail. Markup has to describe what is visibly on the page. Question markup with no visible questions, an author who appears nowhere, a rating pulled from a different site: each of those is a policy breach rather than a clever shortcut, and Google states plainly that markup which does not match the visible content can cost you the eligibility you were chasing.
The newer reason
What schema markup does for AI answers
The rich-result argument for markup has been shrinking for years. The reason to keep writing it has changed, and it is now the stronger of the two.
An answer engine composing a reply has to decide what a page is, who stands behind it and whether the two agree with what it already believes. Clean markup makes that resolution cheap and unambiguous: this organisation, this address, this author, this date, stated the same way across every page of the site. Messy or absent markup does not make the page invisible, it makes it ambiguous, and an ambiguous source is a poor candidate for a sentence a model has to be confident about.
That is why we treat markup as identity work rather than decoration. It is also the least glamorous half of generative engine optimisation, and the half that quietly decides whether the rest of the effort resolves to your brand or to nobody in particular. If you want the reasoning behind the rest of it, start with what generative engine optimisation is and how to get cited by AI in Australia.
Entity consistency is the practical version of all this: the same name, the same address, the same phone number, written identically wherever they appear. It sounds trivial. It is the single most common thing we correct on an Australian site, and it is covered properly in entity SEO.
The failure modes
Five ways schema markup goes wrong
The markup and the page disagree
The most common and the most expensive. Questions in the markup that appear nowhere on the page, an author who does not exist, a date that never updates. Google's guidelines treat this as a breach, not an oversight, and the penalty is losing eligibility across the site rather than on the one page.
A plugin adds everything
Install a plugin, tick every box, and you can end up declaring a page to be four different things at once. An engine resolving contradictions will usually resolve them by ignoring you. Fewer types, stated correctly, beat a full set stated loosely.
Self-awarded ratings
A star rating the business generated about itself is the clearest example of markup that damages the thing it was added to build. If the rating cannot be traced to something a person could verify, leave it out.
It only exists on the homepage
Identity markup that appears once does not establish identity. The organisation block belongs on every page, worded identically, or the consistency it exists to prove is not there.
Nobody checks it again
Markup breaks silently. A template change, a plugin update, a migration, and the block is gone or half-formed with nothing on the page to show it. Whatever else you do, check it on a schedule rather than after a launch.
Google publishes a rich results test and reports structured data errors in Search Console. Both are free, and between them they catch most of the list above. The one they cannot catch is markup that is valid and simply untrue, which is why a person has to read it.
Start here
Not sure what your site is telling the engines?
Get a free audit and we will read your markup the way a search engine and an answer engine do, then tell you what is missing, what contradicts itself and what is quietly costing you nothing but doing nothing.
Get your free auditGood questions
Schema markup FAQs
What is schema markup in simple terms?
It is a block of labelled facts added to a page so software does not have to guess what the page is about. It states things the page already contains, such as the kind of business, the address, the author and the publication date, in a shared vocabulary called schema.org. A visitor never sees it, and it changes nothing about how the page reads.
Does schema markup improve rankings?
Not directly. Google's position has been consistent: structured data helps it understand a page and makes the page eligible for certain result types. Eligibility is not placement. The realistic gains are a better-presented result, fewer misunderstandings about who you are, and cleaner resolution when an AI answer engine has to decide whether to trust the page.
Is FAQ schema still worth adding in 2026?
Not for the rich result. In August 2023 Google restricted frequently-asked-question rich results to well-known, authoritative government and health websites, and they have not been shown regularly for other sites since. If your page genuinely has visible questions and answers, the markup does no harm and helps a machine parse them, but nobody should be paying for it as a way to win space on the results page.
What happened to HowTo schema?
It was withdrawn. Google limited how-to rich results to desktop in August 2023, removed them from desktop in September 2023, and deleted the supporting documentation. The search appearance and the rich result report were retired with it. Advice that still recommends it predates the change.
Which schema types should an Australian business actually use?
Organisation on every page, local business if you have an address or a service area, article on anything you publish, and breadcrumb so your position on the site is explicit. Add product if you sell things. Those five cover nearly all of the available value; everything beyond them should be justified by a documented search feature or a genuine parsing benefit.
Should I use JSON-LD, microdata or RDFa?
JSON-LD, in almost every case. It sits in the page source as one self-contained block, so it survives a redesign that would break markup woven through the visible HTML. It is also the format Google recommends. Microdata and RDFa still work, but they tie your data to your layout for no benefit.
Can schema markup get you penalised?
Markup that does not match what is visible on the page breaches Google's structured data guidelines, and the consequence is losing rich result eligibility, potentially across the whole site rather than the one page. Invented ratings, questions that appear nowhere and authors who do not exist are the usual causes. Valid markup that is simply untrue is the hardest kind to catch, because no automated test will flag it.
Keep rolling