Crawling is how search engines discover pages: automated programs called crawlers download the text, images and video on the pages they find. Google's crawler is called Googlebot.
Why it matters. Crawlers find new pages by following links and reading sitemaps, so a page that nothing links to may never be found.
Sources: Google: In-depth guide to how Google Search works, Google: Link best practices for Google
Indexing is the stage where a search engine analyses a page's text, images and video and stores what it learns in its index, a very large database. Only indexed pages can appear in its results.
Why it matters. Google does not guarantee it will index a page, and a page must be indexed to appear as a link in AI Overviews or AI Mode.
Sources: Google: In-depth guide to how Google Search works, Google: AI features and your website
Crawl budget is the set of URLs on a site that Google can and wants to crawl. It combines how much crawling the server can handle with how much Google wants to crawl.
Why it matters. Google's guide is written for large sites, such as those with over a million pages changing weekly. If your new pages are crawled the day they are published, Google says you do not need its guide.
Source: Google: Crawl budget management for large sites
robots.txt is a file at the root of a site that tells search engine crawlers which URLs they may access. It is mainly used to avoid overloading a site with requests.
Why it matters. It does not keep a page out of Google, and a blocked page can still be indexed if other sites link to it. Use noindex or a password for that.
Sources: Google: Introduction to robots.txt, Google: How to write and submit a robots.txt file, Google: Block Search indexing with noindex
An XML sitemap is a file that tells search engines about the pages, videos and other files on a site, and how they relate.
Why it matters. It helps search engines discover URLs but does not guarantee they are crawled or indexed. Google says a well-linked site of about 500 pages or fewer may not need one.
Source: Google: Learn about sitemaps
A canonical URL is the address Google chooses to represent a set of duplicate or very similar pages. A site can state its preference with rel="canonical", which Google treats as a hint, not a rule.
Why it matters. Google usually shows the canonical page in results and crawls it most often, so make the version you want found the clear choice.
Source: Google: What is URL canonicalisation
A 301 redirect is a server response that tells browsers and search engines a page has moved permanently to a new address. Google follows it and uses it as a signal that the new address should be the canonical one.
Why it matters. When URLs change in a redesign or migration, Google recommends permanent server-side redirects wherever possible, so visitors and search engines reach the right page.
Source: Google: Redirects and Google Search
Structured data is code, such as JSON-LD, that describes a page's content in a standard format, such as who published it or what it offers, so search engines can classify it.
Why it matters. It can make a page eligible for rich results, with no guarantee, and Google says it is not required for its generative AI features. Mark up only what readers can see.
Sources: Google: Introduction to structured data markup, Google: General structured data guidelines, Google: Guide to optimising for generative AI search
IndexNow is a way for a website to tell participating search engines straight away when a URL is added, updated or deleted. A submission to one engine is shared with the others.
Why it matters. Bing, Naver, Seznam.cz, Yandex and Yep take part, Google does not, and a submission does not guarantee indexing.
Sources: IndexNow: IndexNow and its participating search engines, IndexNow: IndexNow frequently asked questions