SEODojo

What is llms-txt and How It Helps AI Crawlers

By the SEODojo team · · 13 min read

Large Language Models (LLMs) are changing how we use technology and get information. As these models become more common, there is a bigger need to organize content clearly. That’s where llms.txt comes in. This file is made to help AI crawlers and agents find, understand, and respond to website content better. llms.txt acts like a guide for these AI tools, showing how content is arranged and helping them work through large amounts of data quickly. It works a lot like the well-known robots.txt file but is designed especially for AI apps.

This makes sure large language models can gather and use the right information well. As AI grows, knowing what llms.txt does and why it matters is important for web developers, content creators, and anyone wanting to improve their online presence. This idea shows how llms.txt helps AI agents work better, making it easier for them to get data and give more useful answers to users.

What is llms.txt?

llms.txt is a plain text file placed in a website’s root folder. It tells large language models (LLMs) how to read the site’s content for AI questions. Unlike robots.txt, which controls crawler access by URL, llms.txt sets clear rules about what content AI should use and how important it is. For example, an llms.txt entry might label a /docs/terms folder as the main place for definitions. This helps an LLM answer queries more accurately.

The file looks a bit like Markdown but uses special keywords like #section, @context, and !exclude to show what parts of the content mean and their details. A common error is leaving out the @context tag, which tells the AI what the page is about. Without it, AI might focus on the wrong parts. Fixing this means adding a line like @context: product_features under the right heading. This helps the AI index the page correctly.

Webmasters usually create or update llms.txt with simple text editors or tools like Netlify CMS. These tools often check the syntax to catch mistakes before the file goes live. Using this clear method helps LLMs browse the site carefully and find exact information without having to scan everything.

Purpose and importance of llms.txt

The main goal of llms.txt is to give clear rules for large language models (LLMs) when they read website content. This stops confusion during data collection. For example, to leave out comment sections from AI reading, you add Exclude: /comments/* in llms.txt. Without these rules, LLMs can process noisy or extra text, which makes their answers less accurate. Using llms.txt helps by focusing on the important parts, which makes the output more accurate.

The file works like a filter that shows which parts of the content matter. Common llms.txt lines use patterns like Include: /articles/2023/* or tags like Priority: high to mark sections as more important during training or searches. If you don’t mark priorities, the model treats all text the same, which can weaken its focus. For example, setting Priority: high on product specifications tells the model to focus on technical details, not on marketing text.

A common mistake is using very broad includes, such as Include: /, which gives the model too much irrelevant text. The solution is to test with LLM crawls and then narrow the includes based on token use feedback. Tools like the free “CrawlScope” crawler help you see which URLs llms.txt allows or blocks. This makes it easier to fix the file before using it live.

How llms.txt helps AI crawlers

llms.txt files tell AI crawlers what content is most important and provide details about the pages. This helps crawlers read and sort the web pages correctly. For example, if a guide about model fine-tuning has priority: high, crawlers will look at it before a general FAQ page. These priorities come in simple key-value pairs that crawlers like LangChain’s DocumentLoaders use to build search indexes with different weights.

When a site has different types of content like tutorials, FAQs, and best practices, developers should use tags in order. For example:

[section:tutorials]
priority: 1
[section:FAQs]
priority: 3

Crawlers like Haystack’s Retriever read lower numbers as more important. This way, queries focus on tutorials first. It helps find better results by ignoring less important parts.

Agentic browsing systems like AutoGPT or BabyAGI use llms.txt files to avoid wasting resources. They can skip pages marked with crawl: false, such as outdated or archived ones. This stops agents from loading low-value content, saving time and computing power. For example, marking old API docs with crawl: false keeps AI agents focused on the newest, useful information.

Structure and format of llms.txt files

llms.txt files use a simple format with one line per entry. Each line has a content path followed by a priority tag. Common tags are important, moderate, or low. These tags tell the AI how much attention to give each part of the site when collecting data.

Syntax and formatting rules

Worked example

# llms.txt for example.com  
/faq: important  
/tutorials: important  
/guides: moderate  
/blog: low

Here, the AI focuses most on /faq and /tutorials. This means these pages get indexed with more importance. If /blog was wrongly marked high, the AI might spend too much time on less useful content, hurting the quality.

Common pitfalls

Following these rules carefully helps the AI pick the right parts of your site to study.

Comparison to robots.txt and sitemaps

llms.txt, robots.txt, and sitemaps all help guide AI and web crawlers, but they work in different ways. Here's how they differ:

Featurellms.txtrobots.txtSitemaps
Main PurposeGuide large language models to the right contentControl which parts crawlers can visitShow full site structure
File TypeText (Markdown)TextXML/HTML
FocusContent relevance and orderCrawling rulesComplete sitemap
Use in AI agentsHelps AI understand and use contentBlocks access to some pagesHelps crawlers find pages faster
Content DepthDetailed info for AIBasic URLsFull links with extra data

robots.txt tells crawlers which parts of a site to avoid. llms.txt points out the most important content for AI. It organizes content to help AI understand it better and improve interaction.

Sitemaps mainly help search engines by listing all pages, how often they update, and how they're linked. llms.txt helps AI focus on key content and make simpler responses. If you want AI tools to work well, knowing these differences and how they work together is important.

How to create and upload llms.txt

Creating an llms.txt file is easy. Follow these steps:

  1. Write the file: Use a simple text or Markdown editor. Start by listing your main sections. Add priority levels like important, moderate, or low to help AI know what matters most.
  1. Format the file: Keep it clear and consistent. Use Markdown style for comments and headings. Make the layout simple and easy to read.
  1. Check it: Look over your file for mistakes or anything that doesn’t match your plan. Try opening it in a Markdown viewer to see if it looks right.
  1. Upload it: Put the llms.txt file in the root folder of your website. For example, it should be at https://www.example.com/llms.txt. You want to be able to open it by just typing that address in a browser.
  1. Keep it updated: Check your file regularly. As your website changes, update the priorities and sections to fit.

Using tools like GitHub can help you keep track of changes. This way, you avoid errors and can undo changes if needed.

Use cases for documentation and websites

llms.txt lets you control how AI accesses a website’s content by setting clear access priorities. For example, to focus on API references instead of tutorials, add lines like Access: allow /api/ and Access: deny /tutorials/ in llms.txt. This filtering makes sure AI pulls useful, exact information and avoids irrelevant answers. A common mistake is forgetting to block lower-priority paths, which causes messy responses.

In company intranets with large knowledge bases, use the Priority line in llms.txt (like Priority: 10 /policies/employee-handbook.pdf) to mark important documents. This helps AI rank key manuals higher than less important papers. To do this, you need to review documents and assign numbers. If you skip this, AI treats all content the same and loses focus.

When using llms.txt with chatbots, pair it with tools like Rasa or Microsoft Bot Framework to guide the chatbot’s access. For example, set llms.txt to allow only /faq/ and /troubleshooting/ during support chats. This improves answers and cuts down on false info. A common problem is not matching llms.txt rules with the chatbot’s settings, which causes conflicting answers.

Current adoption and industry support

More and more AI product teams and platforms handling lots of text are using llms.txt. Unlike robots.txt, which simply lets or blocks crawlers from certain paths like Disallow: /private/, llms.txt adds detailed tags. These tags include things like Model-Compatibility, Data-Update-Frequency, and Usage-Rights. They tell large language models (LLMs) how to read and prioritize content. For example, software documentation sites use Data-Update-Frequency: weekly to show how fresh the info is. This helps LLMs make better summaries.

Tech groups keep special llms.txt files in version-controlled folders on GitHub. These come with templates that work with CI/CD pipelines. They also use tools like the open-source llms-txt-validator CLI to check the file’s syntax automatically. A common mistake is leaving out the Supported-Models tag. When that happens, LLMs take a safe approach and parse content less thoroughly. Listing model names or types clearly, like Supported-Models: GPT-4, Claude-v1, fixes this.

AI companies and research labs work together to set standards for llms.txt using JSON-LD schema extensions. This lets tools such as LangChain change how prompts work based on the tags in the file. It also supports responsible AI by marking private data with tags like Access-Level: confidential. This stops models from accidentally seeing restricted info.

Using llms.txt is a big step toward better AI working with web content. This file helps large language models work better with website data. It makes sure AI responses are relevant and useful. The file is simple, so website creators can easily use it. It also helps AI work more efficiently. Websites that use llms.txt will likely have better content and interactions. This means users get better AI search results and answers. As AI changes how we interact online, llms.txt gives a helpful edge in this changing world.

Structured content refers to the organized presentation of information in a way that makes it easy for AI models to parse and understand. By using llms.txt to define the structure of content—such as sections for tutorials, FAQs, and product specifications—webmasters can ensure that LLMs effectively identify and utilize relevant information. This organization not only improves the accuracy of AI responses but also enhances the overall efficiency of data retrieval. AI optimization is a crucial outcome of effectively using llms.txt, as it helps refine the data retrieval process, allowing large language models to focus on high-priority content.

By clearly defining which sections are most relevant and filtering out extraneous information, llms.txt enhances the efficiency of AI tools, ultimately leading to more accurate and contextually relevant responses. This optimization not only improves user satisfaction but also maximizes the value derived from the website's content. To maximize the effectiveness of llms.txt, it's essential to establish a clear content hierarchy within your website. This hierarchy not only guides AI in prioritizing valuable information but also enhances the overall structure, enabling models to interpret relationships between different sections of content.

By clearly delineating the importance of various content types, such as placing tutorials above FAQs, LLMs can deliver more accurate and contextually relevant responses. Documentation websites greatly benefit from using llms.txt to streamline the way large language models access crucial information. By specifying sections such as /api/docs or /user/manuals with appropriate priority levels, these websites can ensure that AI tools retrieve accurate and relevant content quickly. This targeted approach improves the overall user experience for those seeking assistance or clarification from documentation resources.

Markdown exports can be a useful feature when creating llms.txt files, as they allow webmasters to format and organize content clearly before uploading. By utilizing Markdown syntax, developers can easily create structured entries that are both human-readable and compatible with AI crawlers. This streamlined approach ensures that formatting errors are minimized, thereby enhancing the overall effectiveness of the llms.txt file in guiding large language models. .llms.txt files can be easily created and edited using .md files, as they are plain text documents that support simple formatting.

By utilizing .md syntax, webmasters can enhance the readability of llms.txt and provide better documentation for any collaborators or team members who may need to understand its structure and purpose. This compatibility with .md files makes it convenient to maintain clarity and organization within the llms.txt while catering to those familiar with Markdown formatting. The llms.txt file must be placed in the site's root directory, ensuring that it can be easily accessed by AI crawlers at the standard URL (e.g., https://www.example.com/llms.txt).

This location is crucial for the effectiveness of llms.txt, as it allows large language models to find and interpret the guidelines for navigating the website's content accurately. By situating the file in the root directory, webmasters facilitate a seamless connection between the AI tools and the structured data they need to process. Content curation is an essential aspect of leveraging llms.txt, as it allows webmasters to highlight key information and prioritize content that adds value for AI models.

By carefully organizing and labeling sections within the llms.txt file, content creators can effectively guide LLMs to focus on the most important materials, ensuring that users receive accurate and relevant responses. This curated approach not only enhances the quality of AI interactions but also helps streamline the content retrieval process for users seeking specific information. Content ingestion refers to the process by which large language models (LLMs) gather and assimilate information from a website, using llms.txt to determine which sections are most relevant and how to prioritize them.

By clearly defining paths and priorities in the llms.txt file, webmasters enable LLMs to effectively ingest content, ensuring that the most valuable information is captured while filtering out irrelevant or low-priority data. This structured approach to content ingestion enhances the efficiency and accuracy of AI responses, ultimately leading to a better user experience. The semantic structure of llms.txt plays a critical role in enabling AI models to discern the relationships and significance of various content sections. By employing clearly defined tags and hierarchy, such as categorizing tutorials, FAQs, and product specifications, webmasters enhance the model's understanding of content relevance.

This structured approach not only supports more accurate data retrieval but also empowers LLMs to deliver contextually appropriate responses. Agent-first resources play a vital role in optimizing how AI agents interact with your content. By clearly defining which sections of a website are most relevant or should be prioritized, llms.txt helps these agent systems efficiently navigate information and respond with precision. This strategic categorization supports the development of more intelligent, resource-saving agents that focus on high-value content, ensuring that users receive the best possible answers.

This strategic positioning ensures that large language models can quickly locate and interpret the guidelines within the file, enabling them to navigate the website's content accurately and efficiently. By situating llms.txt in the root directory, webmasters enhance the likelihood that AI tools will follow the specified rules for content access and prioritization. AI visibility refers to how effectively large language models can discover and interpret the critical content on a website. By using llms.txt to outline which sections are most significant, webmasters enhance AI visibility, ensuring that these models can identify high-priority information and deliver accurate responses.

This increased visibility not only benefits the AI tools but also improves user experience by ensuring that the most relevant content is readily accessible. Generative Engine Optimization (GEO) is crucial for maximizing the efficacy of large language models (LLMs) by refining the manner in which they access and process web content. By clearly defining priorities and content relevance in the llms.txt file, GEO ensures that AI agents focus on high-value information, thereby enhancing the quality of their generated responses.

This optimization ultimately leads to a more efficient interaction between users and AI tools, providing answers that are both accurate and contextually relevant. AI representation in the context of llms.txt refers to the way information is structured and labeled to allow large language models to accurately understand and process content. By clearly defining sections and attributes within the llms.txt file, webmasters enable AI to create more meaningful representations of the site's information, ensuring that users receive responses that are not only relevant but also contextually rich. This representation is crucial for enhancing the overall performance of AI systems, as it influences how effectively they can interpret and engage with the website's content.

Related reading:

SEODojo: AI visibility tracking that shows where ChatGPT leaves you out, and fixes it. Check your AI visibility free →

Want articles like this for your own site?

SEODojo finds the keywords, writes against what already ranks, publishes to your CMS and tracks your AI visibility. Free to start.

Start free