llms-txt/search.json
github-actions[bot] 0ac1627140 deploy: 329488a27b
2026-08-26 22:59:03 +00:00

288 lines
No EOL
50 KiB
JSON
Raw Permalink Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

[
{
"objectID": "core.html",
"href": "core.html",
"title": "Python source",
"section": "",
"text": "The llms.txt file spec is for files located in the path llms.txt of a website (or, optionally, in a subpath). llms-sample.txt is a simple example. A file following the spec contains the following sections as markdown, in the specific order:\n\nAn H1 with the name of the project or site. This is the only required section\nA blockquote with a short summary of the project, containing key information necessary for understanding the rest of the file\nZero or more markdown sections (e.g. paragraphs, lists, etc) of any type, except headings, containing more detailed information about the project and how to interpret the provided files\nZero or more markdown sections delimited by H2 headers, containing “file lists” of URLs where further detail is available\n\nEach “file list” is a markdown list, containing a required markdown hyperlink [name](url), then optionally a : and notes about the file.\n\n\nHere’s the start of a sample llms.txt file we’ll use for testing:\n\nsamp = Path('llms-sample.txt').read_text()\nprint(samp[:480])\n\n# FastHTML\n\n> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore's `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.\n\nRemember:\n\n- Use `serve()` for running uvicorn (`if __name__ == \"__main__\"` is not needed since it's automatic)\n- When a title is needed with a response, use `Titled`; note that that already wraps children in `Container`, and already includes both the meta title as well as the H1 element",
"crumbs": [
"Code",
"Python source"
]
},
{
"objectID": "core.html#introduction",
"href": "core.html#introduction",
"title": "Python source",
"section": "",
"text": "The llms.txt file spec is for files located in the path llms.txt of a website (or, optionally, in a subpath). llms-sample.txt is a simple example. A file following the spec contains the following sections as markdown, in the specific order:\n\nAn H1 with the name of the project or site. This is the only required section\nA blockquote with a short summary of the project, containing key information necessary for understanding the rest of the file\nZero or more markdown sections (e.g. paragraphs, lists, etc) of any type, except headings, containing more detailed information about the project and how to interpret the provided files\nZero or more markdown sections delimited by H2 headers, containing “file lists” of URLs where further detail is available\n\nEach “file list” is a markdown list, containing a required markdown hyperlink [name](url), then optionally a : and notes about the file.\n\n\nHere’s the start of a sample llms.txt file we’ll use for testing:\n\nsamp = Path('llms-sample.txt').read_text()\nprint(samp[:480])\n\n# FastHTML\n\n> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore's `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.\n\nRemember:\n\n- Use `serve()` for running uvicorn (`if __name__ == \"__main__\"` is not needed since it's automatic)\n- When a title is needed with a response, use `Titled`; note that that already wraps children in `Container`, and already includes both the meta title as well as the H1 element",
"crumbs": [
"Code",
"Python source"
]
},
{
"objectID": "core.html#reading",
"href": "core.html#reading",
"title": "Python source",
"section": "Reading",
"text": "Reading\nWe’ll implement parse_llms_file to pull out the sections of llms.txt into a simple data structure.\n\n\n\nsearch\ndef search(\n pat, txt, flags:int=0\n):\nDictionary of matched groups in pat within txt\n\n\n\n\n\nnamed_re\ndef named_re(\n nm, pat\n):\nPattern to match pat in a named capture group\n\n\n\n\n\nopt_re\ndef opt_re(\n s\n):\nPattern to optionally match s\n\n\nWe’ll work “outside in” so we can test the innermost matches as we go.\n\nParse links\n\nlink = '- [FastHTML quick start](https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md): A brief overview of FastHTML features'\n\n\ntitle = named_re('title', r'[^\\]]+')\npat = fr'-\\s*\\[{title}\\]'\nsearch(pat, samp)\n\n{'title': 'FastHTML quick start'}\n\n\n\nurl = named_re('url', r'[^\\)]+')\npat += fr'\\({url}\\)'\nsearch(pat, samp)\n\n{'title': 'FastHTML quick start',\n 'url': 'https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md'}\n\n\n\ndesc = named_re('desc', r'.*')\npat += opt_re(fr':\\s*{desc}')\nsearch(pat, link)\n\n{'title': 'FastHTML quick start',\n 'url': 'https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md',\n 'desc': 'A brief overview of FastHTML features'}\n\n\n\n\n\nparse_link\ndef parse_link(\n txt\n):\nParse a link section from llms.txt\n\n\n\nparse_link(link)\n\n{'title': 'FastHTML quick start',\n 'url': 'https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md',\n 'desc': 'A brief overview of FastHTML features'}\n\n\n\nparse_link('-[foo](http://foo)')\n\n{'title': 'foo', 'url': 'http://foo', 'desc': None}\n\n\n\n\nParse sections\n\nsections = '''First bit.\n\n## S1\n\n-[foo](http://foo)\n- [foo2](http://foo2): stuff\n\n## S2\n\n- [foo3](http://foo3)'''\n\n\nstart,*rest = re.split(fr'^##\\s*(.*?$)', sections, flags=re.MULTILINE)\nstart\n\n'First bit.\\n\\n'\n\n\n\nrest\n\n['S1',\n '\\n\\n-[foo](http://foo)\\n- [foo2](http://foo2): stuff\\n\\n',\n 'S2',\n '\\n\\n- [foo3](http://foo3)']\n\n\n\nd = dict(chunked(rest, 2))\nd\n\n{'S1': '\\n\\n-[foo](http://foo)\\n- [foo2](http://foo2): stuff\\n\\n',\n 'S2': '\\n\\n- [foo3](http://foo3)'}\n\n\n\nlinks = d['S1']\nlinks.strip()\n\n'-[foo](http://foo)\\n- [foo2](http://foo2): stuff'\n\n\n\n_parse_links(links)\n\n[{'title': 'foo', 'url': 'http://foo', 'desc': None},\n {'title': 'foo2', 'url': 'http://foo2', 'desc': 'stuff'}]\n\n\n\nstart, sects = _parse_llms(samp)\nstart\n\n'# FastHTML\\n\\n> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore\\'s `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.\\n\\nRemember:\\n\\n- Use `serve()` for running uvicorn (`if __name__ == \"__main__\"` is not needed since it\\'s automatic)\\n- When a title is needed with a response, use `Titled`; note that that already wraps children in `Container`, and already includes both the meta title as well as the H1 element.'\n\n\n\ntitle = named_re('title', r'.+?$')\nsumm = named_re('summary', '.+?$')\nsumm_pat = opt_re(fr\"^>\\s*{summ}$\")\ninfo = named_re('info', '.*')\n\n\npat = fr'^#\\s*{title}\\n+{summ_pat}\\n+{info}'\nsearch(pat, start, (re.MULTILINE|re.DOTALL))\n\n{'title': 'FastHTML',\n 'summary': 'FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore\\'s `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.',\n 'info': 'Remember:\\n\\n- Use `serve()` for running uvicorn (`if __name__ == \"__main__\"` is not needed since it\\'s automatic)\\n- When a title is needed with a response, use `Titled`; note that that already wraps children in `Container`, and already includes both the meta title as well as the H1 element.'}\n\n\n\n\n\nparse_llms_file\ndef parse_llms_file(\n txt\n):\nParse llms.txt file contents in txt to an AttrDict\n\n\n\nllmsd = parse_llms_file(samp)\nllmsd.summary\n\n'FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore\\'s `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.'\n\n\n\nllmsd.sections.Examples\n\n[{'title': 'Todo list application', 'url': 'https://raw.githubusercontent.com/AnswerDotAI/fasthtml/main/examples/adv_app.py', 'desc': 'Detailed walk-thru of a complete CRUD app in FastHTML showing idiomatic use of FastHTML and HTMX patterns.'}]",
"crumbs": [
"Code",
"Python source"
]
},
{
"objectID": "core.html#xml-conversion",
"href": "core.html#xml-conversion",
"title": "Python source",
"section": "XML conversion",
"text": "XML conversion\nFor some LLMs such as Claude, XML format is preferred, so we’ll provide a function to create that format.\n\n\n\nget_doc_content\ndef get_doc_content(\n url\n):\nFetch content from local file if in nbdev repo.\n\n\n\n\n\nmk_ctx\ndef mk_ctx(\n d, optional:bool=True, n_workers:NoneType=None\n):\nCreate a Project with a Section for each H2 part in d, optionally skipping the ‘optional’ section.\n\n\n\nctx = mk_ctx(llmsd)\nprint(to_xml(ctx, do_escape=False)[:260]+'...')\n\n{'title': 'FastHTML quick start', 'url': 'https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md', 'desc': 'A brief overview of FastHTML features'}{'title': 'HTMX reference', 'url': 'https://raw.githubusercontent.com/bigskysoftware/htmx/master/www/content/reference.md', 'desc': 'Brief description of all HTMX attributes, CSS classes, headers, events, extensions, js lib methods, and config options'}\n\n{'title': 'Starlette quick guide', 'url': 'https://gist.githubusercontent.com/jph00/e91192e9bdc1640f5421ce3c904f2efb/raw/61a2774912414029edaf1a55b506f0e283b93c46/starlette-quick.md', 'desc': {}}\n{'title': 'Todo list application', 'url': 'https://raw.githubusercontent.com/AnswerDotAI/fasthtml/main/examples/adv_app.py', 'desc': 'Detailed walk-thru of a complete CRUD app in FastHTML showing idiomatic use of FastHTML and HTMX patterns.'}\n{'title': 'Starlette full documentation', 'url': 'https://gist.githubusercontent.com/jph00/809e4a4808d4510be0e3dc9565e9cbd3/raw/9b717589ca44cedc8aaf00b2b8cacef922964c0f/starlette-sml.md', 'desc': 'A subset of the Starlette documentation useful for FastHTML development.'}\n<project title=\"FastHTML\" summary='FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore's `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.'>Remember:\n\n- Use `serve()` for running uvic...\n\n\n\n\n\nget_sizes\ndef get_sizes(\n ctx\n):\nGet the size of each section of the LLM context\n\n\n\nget_sizes(ctx)\n\n{'docs': {'FastHTML quick start': 35321,\n 'HTMX reference': 28365,\n 'Starlette quick guide': 7936},\n 'examples': {'Todo list application': 16032},\n 'optional': {'Starlette full documentation': 48331}}\n\n\n\nPath('../fasthtml.md').write_text(to_xml(ctx, do_escape=False))\n\n137125\n\n\n\n\n\ncreate_ctx\ndef create_ctx(\n txt, optional:bool=False, n_workers:NoneType=None\n):\nA Project with a Section for each H2 part in txt, optionally skipping the ‘optional’ section.\n\n\n\n\n\nllms_txt2ctx\ndef llms_txt2ctx(\n fname:str, # File name to read\n optional:bool_arg=False, # Include 'optional' section?\n n_workers:int=None, # Number of threads to use for parallel downloading\n save_nbdev_fname:str=None, # save output to nbdev `{docs_path}` instead of emitting to stdout\n):\nPrint a Project with a Section for each H2 part in file read from fname, optionally skipping the ‘optional’ section.\n\n\n\n!llms_txt2ctx llms-sample.txt > ../fasthtml.md",
"crumbs": [
"Code",
"Python source"
]
},
{
"objectID": "ed.html",
"href": "ed.html",
"title": "ed, the standard text editor",
"section": "",
"text": "ed, the standard text editor\nIn order to understand how llms.txt can be used with editors and IDEs, let’s look at how ed, the standard text editor, could work (assuming it’s updated to use this proposal). In our example we will look at how the user might then tell ed to retrieve the LLM docs from fastht.ml/docs, and then use the results to write a simple FastHTML web app.\nEven if you use a non-standard editor or IDE such as vscode, Cursor, vim, or Emacs, your software’s interaction with /llms.txt would look similar to this general approach.\n$ ed\n* H\nOur user starts ed and enables helpful error messages (just for the purpose of this walkthru - obviously a real ed user doesn’t need “helpful error messages”).\n* L fastht.ml/docs\nChecking for /llms.txt at fastht.ml/docs...\nFound /llms.txt. Parsing...\nFetching URLs from \"Docs\" section... Fetching URLs from \"Examples\" section...\nSkipping \"Optional\" section for brevity.\nCreating XML-based context for Claude... Context created and loaded.\nThe user invokes the hypothetical L (load) command, which in this LLM-enhanced version of ed retrieves and processes the llms.txt file. ed checks for the file (if it didn’t exist, it would fall back to scraping the HTML of the website the old-fashioned way), parses it, fetches the relevant URLs, and creates an XML-based context suitable for Claude (perhaps an ed config file could be used to choose what LLM to use, and would determine how the context is formatted). All of this happens with the characteristic silence of ed, broken only by these reassuring progress messages.\n* x Create a simple FastHTML app which outputs 'Hello, World!', in a <div>.\nAnalyzing context and prompt...\nGenerating FastHTML app...\nApp written to buffer.\nNext, our user invokes the hypothetical x (eXecute AI) command, providing instructions for the LLM to create a simple FastHTML app. In the world of LLM-enhanced ed, this is understood as a request to generate code based on the given prompt and the previously loaded context.\n* n\n5\n* p\nfrom fasthtml.common import *\napp,rt = fast_app()\n@rt\ndef index(): return div(\"Hello, World!\")\nserve()\nThe editor analyzes the loaded context along with the provided prompt, generates the FastHTML app, and writes it to the buffer. The user then views the generated app line count (n) and contents (p), marveling at how much functionality is packed into those 5 lines.\n*w hello_world.py\n5\n*q\nFinally, our user saves the app to a file and quits ed, presumably to run their new FastHTML app and reflect on the unexpected productivity boost provided by their trusty line editor.",
"crumbs": [
"Editors and IDEs",
"`ed`, the standard text editor"
]
},
{
"objectID": "index.html",
"href": "index.html",
"title": "The /llms.txt file, v2",
"section": "",
"text": "Agents now use websites constantly: a coding agent fetches a library’s documentation to get an API call right, and a chat assistant with search reads pages to answer questions about a product. When this proposal was first written in 2024, this was largely a prediction. Today it is routine.\nBut web pages are built for people. An HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise. Context windows, while larger than they were, are still too small for most websites in their entirety, and every wasted token costs time and money. Agents are best served by concise, expert-level information gathered in a single, accessible location.\nThis is v2 of the proposal, updated based on what I learned from two years of adoption: thousands of sites publish an llms.txt file, documentation platforms generate one automatically, and Chrome’s Lighthouse audits sites for one as part of its agentic browsing checks. The AI labs themselves publish llms.txt files for their own developer docs: OpenAI, Anthropic, and Gemini. The Changes page describes what changed since v1, and why.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#background",
"href": "index.html#background",
"title": "The /llms.txt file, v2",
"section": "",
"text": "Agents now use websites constantly: a coding agent fetches a library’s documentation to get an API call right, and a chat assistant with search reads pages to answer questions about a product. When this proposal was first written in 2024, this was largely a prediction. Today it is routine.\nBut web pages are built for people. An HTML page wraps its information in navigation, ads, and JavaScript, and converting it back into clean text is difficult and imprecise. Context windows, while larger than they were, are still too small for most websites in their entirety, and every wasted token costs time and money. Agents are best served by concise, expert-level information gathered in a single, accessible location.\nThis is v2 of the proposal, updated based on what I learned from two years of adoption: thousands of sites publish an llms.txt file, documentation platforms generate one automatically, and Chrome’s Lighthouse audits sites for one as part of its agentic browsing checks. The AI labs themselves publish llms.txt files for their own developer docs: OpenAI, Anthropic, and Gemini. The Changes page describes what changed since v1, and why.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#proposal",
"href": "index.html#proposal",
"title": "The /llms.txt file, v2",
"section": "Proposal",
"text": "Proposal\n\n\n\nllms.txt logo\n\n\nWe propose adding a /llms.txt markdown file to websites to provide LLM-friendly content. The file can be placed at the site root, or at any path within it, covering the pages under that path. This file offers brief background information, guidance, and links to detailed markdown files.\nllms.txt markdown is human and LLM readable, but is also in a precise format allowing fixed processing methods (i.e. classical programming techniques such as parsers and regex).\nWe furthermore propose that pages with information that agents might need provide a clean markdown version of those pages at the same URL as the original page, either with .md appended (page.html.md) or with the extension replaced by .md (page.md). (URLs without file names should append index.html.md or index.md instead.)\nTo help clients find these files, we recommend using standard link relations: rel=\"alternate\" type=\"text/markdown\" points to the markdown version of a page, and rel=\"describedby\" points to the llms.txt file that covers it. (An llms.txt file describes all pages under its path, so /docs/llms.txt covers everything in /docs/.) These links can be provided as HTML <link> elements, or as an HTTP Link: response header. The header form also works for non-HTML resources, such as the markdown files themselves, and can be added in web server or CDN configuration without modifying any pages. For example:\nLink: </docs/page.html.md>; rel=\"alternate\"; type=\"text/markdown\", </docs/llms.txt>; rel=\"describedby\"\nThe FastHTML project follows these two proposals for its documentation. For instance, here is the FastHTML docs llms.txt, placed at /docs/ to cover just the documentation pages. And here is an example of a regular HTML docs page, along with exact same URL but with a .md extension.\nAgents are expected to view or search llms.txt to find the information they need, then follow the relevant links. The links in an llms.txt file should therefore point to LLM-friendly content, such as the markdown versions of pages described above. The file itself stays small enough to fit in context. The detail lives behind the links, and is fetched only when needed.\nllms.txt files are used most heavily for software documentation, where coding agents follow them to find API references and tutorials. The same structure works anywhere agents need a guided path into a site’s content: a business outlining its structure and policies, a personal site answering questions about someone’s CV, or a school providing access to course information.\nNote that all nbdev projects now create .md versions of all pages by default. All Answer.AI and fast.ai software projects using nbdev have had their docs regenerated with this feature. For an example, see the markdown version of fastcore’s docments module.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#format",
"href": "index.html#format",
"title": "The /llms.txt file, v2",
"section": "Format",
"text": "Format\nAt the moment the most widely and easily understood format for language models is Markdown. Simply showing where key Markdown files can be found is a great first step. Providing some basic structure helps a language model to find where the information it needs can come from.\nThe llms.txt file is unusual in that it uses Markdown to structure the information rather than a classic structured format such as XML. The reason for this is that we expect many of these files to be read by language models and agents. Having said that, the information in llms.txt follows a specific format and can be read using standard programmatic-based tools.\nThe llms.txt file spec is for files named llms.txt, at the root path /llms.txt of a website or at any subpath (e.g. /docs/llms.txt). A file covers the URLs under its path, and where more than one file applies, agents should use the most specific one. A file following the spec contains the following sections as markdown, in the specific order:\n\nAn optional byte-order mark (BOM)\nAn H1 with the name of the project or site. This is the only required section\nA blockquote with a short summary of the project, containing key information necessary for understanding the rest of the file\nZero or more markdown sections (e.g. paragraphs, lists, etc) of any type except headings, containing more detailed information about the project and how to interpret the provided files\nZero or more markdown sections delimited by H2 headers, containing “file lists” of URLs where further detail is available\n\nEach “file list” is a markdown list, containing a required markdown hyperlink [name](url), then optionally a : and notes about the file.\n\n\nHere is a mock example:\n# Title\n\n> Optional description goes here\n\nOptional details go here\n\n## Section name\n\n- [Link title](https://link_url): Optional link details\n\n## Optional\n\n- [Link title](https://link_url)\nThe “Optional” section is used, by convention, for secondary information: links an agent can skip when a shorter context is needed.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#existing-standards",
"href": "index.html#existing-standards",
"title": "The /llms.txt file, v2",
"section": "Existing standards",
"text": "Existing standards\nllms.txt is designed to coexist with current web standards. While sitemaps list all pages for search engines, llms.txt offers a curated overview for LLMs. It can complement robots.txt by providing context for allowed content. The file can also reference structured data markup used on the site, helping LLMs understand how to interpret this information in context.\nThe approach of using a standard filename follows /robots.txt and /sitemap.xml at the site root. And just as index.html gives any path a conventional location for its human-readable entry point, llms.txt gives any path a conventional location for its LLM-readable overview. robots.txt and llms.txt have different purposes. robots.txt lets automated tools know what access to a site is considered acceptable, such as for search indexing bots. llms.txt information is instead used on demand, when an agent needs information about a topic while assisting a user. Our expectation was that llms.txt would mainly be useful for inference rather than training, and that is how it has been used, though training runs could take advantage of the information too.\nAn alternative would be the Well-Known URIs standard (RFC 8615), which reserves the /.well-known/ prefix for metadata files like this one. But well-known URIs exist only at the origin root, and many authors control only a path on a shared host: a GitHub Pages project site, for example, can publish files in its own directory but can never add one to the host’s /.well-known/. Like index.html, an llms.txt describes the path where it sits, something a single root location cannot express. And anyone who can publish content at a path can provide one.\nsitemap.xml is a list of all the indexable human-readable information available on a site. This isn’t a substitute for llms.txt since it:\n\nOften won’t have the LLM-readable versions of pages listed\nDoesn’t include URLs to external sites, even though they might be helpful to understand the information\nWill generally cover documents that in aggregate will be too large to fit in an LLM context window, and will include a lot of information that isn’t necessary to understand the site.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#example",
"href": "index.html#example",
"title": "The /llms.txt file, v2",
"section": "Example",
"text": "Example\nHere’s an example of llms.txt, in this case a cut down version of the file used for the FastHTML project (see also the full version):\n# FastHTML\n\n> FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore's `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.\n\nImportant notes:\n\n- Although parts of its API are inspired by FastAPI, it is *not* compatible with FastAPI syntax and is not targeted at creating API services\n- FastHTML is compatible with JS-native web components and any vanilla JS library, but not with React, Vue, or Svelte.\n\n## Docs\n\n- [FastHTML quick start](https://fastht.ml/docs/tutorials/quickstart_for_web_devs.html.md): A brief overview of many FastHTML features\n- [HTMX reference](https://github.com/bigskysoftware/htmx/blob/master/www/content/reference.md): Brief description of all HTMX attributes, CSS classes, headers, events, extensions, js lib methods, and config options\n\n## Examples\n\n- [Todo list application](https://github.com/AnswerDotAI/fasthtml/blob/main/examples/adv_app.py): Detailed walk-thru of a complete CRUD app in FastHTML showing idiomatic use of FastHTML and HTMX patterns.\n\n## Optional\n\n- [Starlette full documentation](https://gist.githubusercontent.com/jph00/809e4a4808d4510be0e3dc9565e9cbd3/raw/9b717589ca44cedc8aaf00b2b8cacef922964c0f/starlette-sml.md): A subset of the Starlette documentation useful for FastHTML development. \nTo create effective llms.txt files, consider these guidelines:\n\nUse concise, clear language.\nWhen linking to resources, include brief, informative descriptions.\nAvoid ambiguous terms or unexplained jargon.\nTest your file by asking an agent questions about your content, giving it only your llms.txt as a starting point.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#directories",
"href": "index.html#directories",
"title": "The /llms.txt file, v2",
"section": "Directories",
"text": "Directories\nHere are a few directories that list the llms.txt files available on the web:\n\nllmstxt.site\ndirectory.llmstxt.cloud\nllmstxthub.com",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#integrations",
"href": "index.html#integrations",
"title": "The /llms.txt file, v2",
"section": "Integrations",
"text": "Integrations\nMany documentation platforms and CMSs can generate an llms.txt file automatically:\n\nMintlify - Docs platform that generates llms.txt and markdown page versions for every site it hosts\nGitBook - Serves an llms.txt file for published docs sites\nYoast SEO - WordPress plugin that generates and maintains an llms.txt file\nAIOSEO - WordPress plugin with an llms.txt generator\nWix - Generates an llms.txt file for every Wix site\n\nAnd various libraries and plugins are available to integrate the llms.txt specification into your workflow:\n\nJavaScript Implementation - Sample JavaScript implementation\nvitepress-plugin-llms - VitePress plugin that automatically generates LLM-friendly documentation for the website following the llms.txt specification\ndocusaurus-plugin-llms - Docusaurus plugin for generating LLM-friendly documentation following the llmtxt.org standard\nDrupal LLM Support - A Drupal Recipe providing full support for the llms.txt proposal on any Drupal 10.3+ site\nllms-txt-php - A library for writing and reading llms.txt Markdown files\nVS Code PagePilot Extension - PagePilot is a VS Code Chat participant that automatically loads external context (documentation, APIs, README files) to provide enhanced responses.\nserver-llm-txt - MCP server that lets agents fetch and search llms.txt files",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "index.html#next-steps",
"href": "index.html#next-steps",
"title": "The /llms.txt file, v2",
"section": "Next steps",
"text": "Next steps\nThe llms.txt specification is open for community input. A GitHub repository hosts this informal overview, allowing for version control and public discussion. A community discord channel is available for sharing implementation experiences and discussing best practices.",
"crumbs": [
"The /llms.txt file, v2"
]
},
{
"objectID": "nbdev.html",
"href": "nbdev.html",
"title": "How to help LLMs understand your nbdev project",
"section": "",
"text": "This tutorial demonstrates how to add llms.txt to your nbdev project, creating a clear interface between your code and the agents that use it. nbdev already publishes markdown versions of every documentation page by default, so the resources your llms.txt links to are ready-made.\nWhile this guide focuses on nbdev, the underlying principles and tools are framework-agnostic and can help make any codebase more accessible to agents.\nLet’s explore how to implement this.",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "nbdev.html#overview",
"href": "nbdev.html#overview",
"title": "How to help LLMs understand your nbdev project",
"section": "",
"text": "This tutorial demonstrates how to add llms.txt to your nbdev project, creating a clear interface between your code and the agents that use it. nbdev already publishes markdown versions of every documentation page by default, so the resources your llms.txt links to are ready-made.\nWhile this guide focuses on nbdev, the underlying principles and tools are framework-agnostic and can help make any codebase more accessible to agents.\nLet’s explore how to implement this.",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "nbdev.html#the-key-ingredient-llms.txt",
"href": "nbdev.html#the-key-ingredient-llms.txt",
"title": "How to help LLMs understand your nbdev project",
"section": "The key ingredient: llms.txt",
"text": "The key ingredient: llms.txt\nThe foundation of LLM-friendly documentation is the llms.txt file. At its core, it is just a Markdown file with information about your library found at a specific URL (root of your site followed by /llms.txt).\nHowever, it needs to follow a certain structure as outlined in the llms.txt format.\nDo not be intimidated by the specification, though. In reality, it offers a lot of flexibility and by conforming to it you’ll gain access to several very helpful tools that we will look at in a second.\nFirst, let’s start working on our llms.txt file. If you would like to, you can open your favorite editor and start working on an llms.txt for your library as we go along.\nHere is how the llms.txt file could begin:\n# FastHTML\n\n> FastHTML is a python library which...\n\nWhen writing FastHTML apps remember to:\n\n- Thing to remember\nThe required elements are:\n\nthe H1 header (FastHTML)\na blockquote with a short summary of the project (FastHTML is a python library which…)\n\nAnd they can optionally be followed by zero or more paragraphs and lists. Usually, this is the place where you would add a short description of your library.\nThe description can be as simple as this (this is an excerpt from the llms.txt for fastcore):\nHere are some tips on using fastcore:\n\n- **Liberal imports**: Utilize `from fastcore.module import *` freely. The library is designed for safe wildcard imports.\n- **Enhanced list operations**: Substitute `list` with `L`. This provides advanced indexing, method chaining, and additional functionality while maintaining list-like behavior.\n- **Extend existing classes**: Apply the `@patch` decorator to add methods to classes, including built-ins, without subclassing. This enables more flexible code organization.\nBelow are a few ideas on how to make writing the description feel even more seamless:\n\nConsider the content you already have that can be used as a starting point (e.g. your project’s README, blog posts and articles, social media discussions, etc.)\nThink of how you would describe your library to a new team member — this often yields the right balance of precision and comprehension.\nUse an LLM to help you synthetize content from multiple sources into cohesive prose (though you might need to do some post-processing to combat the LLM’s tendency to be verbose).\n\n\nAdding resource sections\nAfter the optional description, you can include zero or more sections starting with an H2 heading and containing links to supplementary resources.\nMarkdown files are strongly recommended here as they offer a good balance of structure and readability. You could attempt linking to other formats, but your results may vary. For instance, HTML tends to be verbose, and formats like CSV rarely contain information that lends itself well to documenting functionality.\nHere’s an example of what this section might look like:\n## Docs\n\n- [Surreal](https://host/README.md): Tiny jQuery alternative with Locality of Behavior\n- [FastHTML quick start](https://host/quickstart.html.md): An overview of FastHTML features\n\n## Examples\n\n- [Todo app](https://host/adv_app.py)\n\n\nThe Optional section\nIf you’d like to, you can include a section with Optional as the heading. Use it for secondary resources: things an agent only needs for a deeper dive, and can skip when context is tight.\nHere is a small example of the Optional section:\n## Optional\n\n- [Starlette docs](https://host/starlette-sml.md): A subset of the Starlette docs\nYour llms.txt file is now complete! Time to give yourself a pat on the back for a job well done and let’s move on to the next, automated step.",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "nbdev.html#enhancing-context-with-api-reference",
"href": "nbdev.html#enhancing-context-with-api-reference",
"title": "How to help LLMs understand your nbdev project",
"section": "Enhancing context with API reference",
"text": "Enhancing context with API reference\nWhile LLMs generally understand high-level concepts, they often struggle with implementation details, especially when their training data is outdated. Providing a comprehensive list of your library’s symbols - functions, classes, and their documentation - helps bridge this gap.\nThis is where the pysym2md library enters the picture. It generates a complete API reference in Markdown, extracting existing docstrings along the way.\nThis is a short excerpt from the apilist.txt for fastcore:\n# fastcore Module Documentation\n\n## fastcore.ansi\n\n> Filters for processing ANSI colors.\n\n- `def strip_ansi(source)`\n Remove ANSI escape codes from text.\n\n- `def ansi2html(text)`\n Convert ANSI colors to HTML colors.\n\n- `def ansi2latex(text)`\n Convert ANSI colors to LaTeX colors.\nThe tool works great even with larger libraries. For instance, generating the API reference for numpy requires just one command:\npysym2md numpy\nTo implement this in your project, generate an apilist.txt, serve it alongside your documentation, and reference it from your llms.txt file.",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "nbdev.html#configuration",
"href": "nbdev.html#configuration",
"title": "How to help LLMs understand your nbdev project",
"section": "Configuration",
"text": "Configuration\nThe final step is to configure your nbdev project to generate and serve these files. This requires three changes:\n\nAdd your llms.txt file to the nbs directory of your project.\nAdd the required dependencies to settings.ini:\n\ndev_requirements = pysym2md\n\nConfigure Quarto’s build process in nbs/_quarto.yml:\n\nproject:\n type: website\n pre-render:\n - pysym2md --output-file apilist.txt nbdev\n resources:\n - \"*.txt\"\nRemember to manually add a link to the generated apilist.txt in your llms.txt file. Once you commit these changes and rebuild your docs, your library will be ready for deeper, more accurate conversations with LLMs!",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "nbdev.html#learning-from-examples",
"href": "nbdev.html#learning-from-examples",
"title": "How to help LLMs understand your nbdev project",
"section": "Learning from examples",
"text": "Learning from examples\nIt is often useful to study how others went about implementing the thing we are working on. The fastcore and FastHTML projects offer a good reference, and you can find additional examples in the llmstxt.site and llmstxt.cloud directories.\nProviding the right context opens up new possibilities for AI-assisted development and exploring topics you might want to learn more about. We hope this guide helps you and the users of your library take advantage of these exciting new tools.",
"crumbs": [
"Tutorials",
"How to help LLMs understand your nbdev project"
]
},
{
"objectID": "domains.html",
"href": "domains.html",
"title": "llms.txt in Different Domains",
"section": "",
"text": "This page has some guidelines and suggestions for how different domains could utilize llms.txt to allow LLMs to better interface with their site if they so choose.\nRemember, when constructing your llms.txt you should “use concise, clear language. When linking to resources, include brief, informative descriptions. Avoid ambiguous terms or unexplained jargon.” Additionally, the best way to determine if your llms.txt works well with LLMs is to test it with them! Here is a minimal way to test Anthropic’s Claude against your llms.txt:\n# /// script\n# requires-python = \">=3.8\"\n# dependencies = [\n# \"claudette\",\n# \"llms-txt\",\n# \"requests\",\n# ]\n# ///\nfrom claudette import *\nfrom llms_txt import create_ctx\n\nimport requests\n\nmodel = models[1] # Sonnet 3.5\nchat = Chat(model, sp=\"\"\"You are a helpful and concise assistant.\"\"\")\n\nurl = 'your_url/llms.txt'\ntext = requests.get(url).text\nllm_ctx = create_ctx(text)\nchat(llm_ctx + '\\n\\nThe above is necessary context for the conversation.')\n\nwhile True:\n msg = input('Your question about the site: ')\n res = chat(msg)\n print('From Claude:', contents(res))\nThe above script utilizes the relatively new uv syntax for python scripts. If you install uv, you can simply run the above script with uv run test_llms_txt.py and it will handle installing the necessary library dependencies in an isolated python environment. Else you can install the requirements manually and run it like any ordinary python script with python test_llms_txt.py.\n\n\nHere is an example llms.txt that a restaurant could construct for consumption by LLMs:\n# Nate the Great's Grill\n\n> Nate the Great's Grill is a popular destination off of Sesame Street that has been serving the community for over 20 years. We offer the best BBQ for a great price.\n\nHere are our weekly hours:\n\n- Monday - Friday: 9am - 9pm\n- Saturday: 11am - 9pm\n- Sunday: Closed\n\n## Menus\n\n- [Lunch Menu](https://host/lunch.html.md): Our lunch menu served from 11am to 4pm every day.\n- [Dinner Menu](https://host/dinner.html.md): Our dinner menu served from 4pm to 9pm every day.\n\n## Optional\n\n- [Dessert Menu](https://host/dessert.md): A subset of the Starlette docs\nAnd here is an example lunch menu taken from Franklin’s BBQ:\n## By The Pound\n\n| Item | Price |\n| -------------- | ----------- |\n| Brisket | 34 |\n| Pork Spare Ribs | 30 |\n| Pulled Pork | 28 |\n\n## Drinks\n\n| Item | Price |\n| -------------- | ----------- |\n| Iced Tea | 3 |\n| Mexican Coke | 3 |\n\n## Sides\n\n| Item | Price |\n| -------------- | ----------- |\n| Potato Salad | 4 |\n| Slaw | 4 |",
"crumbs": [
"Tutorials",
"llms.txt in Different Domains"
]
},
{
"objectID": "domains.html#restaurants",
"href": "domains.html#restaurants",
"title": "llms.txt in Different Domains",
"section": "",
"text": "Here is an example llms.txt that a restaurant could construct for consumption by LLMs:\n# Nate the Great's Grill\n\n> Nate the Great's Grill is a popular destination off of Sesame Street that has been serving the community for over 20 years. We offer the best BBQ for a great price.\n\nHere are our weekly hours:\n\n- Monday - Friday: 9am - 9pm\n- Saturday: 11am - 9pm\n- Sunday: Closed\n\n## Menus\n\n- [Lunch Menu](https://host/lunch.html.md): Our lunch menu served from 11am to 4pm every day.\n- [Dinner Menu](https://host/dinner.html.md): Our dinner menu served from 4pm to 9pm every day.\n\n## Optional\n\n- [Dessert Menu](https://host/dessert.md): A subset of the Starlette docs\nAnd here is an example lunch menu taken from Franklin’s BBQ:\n## By The Pound\n\n| Item | Price |\n| -------------- | ----------- |\n| Brisket | 34 |\n| Pork Spare Ribs | 30 |\n| Pulled Pork | 28 |\n\n## Drinks\n\n| Item | Price |\n| -------------- | ----------- |\n| Iced Tea | 3 |\n| Mexican Coke | 3 |\n\n## Sides\n\n| Item | Price |\n| -------------- | ----------- |\n| Potato Salad | 4 |\n| Slaw | 4 |",
"crumbs": [
"Tutorials",
"llms.txt in Different Domains"
]
},
{
"objectID": "intro.html",
"href": "intro.html",
"title": "Python module & CLI",
"section": "",
"text": "Given an llms.txt file, this provides a CLI and Python API to parse the file and create an XML context file from it. The input file should follow this format:",
"crumbs": [
"Code",
"Python module & CLI"
]
},
{
"objectID": "intro.html#install",
"href": "intro.html#install",
"title": "Python module & CLI",
"section": "Install",
"text": "Install\npip install llms-txt",
"crumbs": [
"Code",
"Python module & CLI"
]
},
{
"objectID": "intro.html#how-to-use",
"href": "intro.html#how-to-use",
"title": "Python module & CLI",
"section": "How to use",
"text": "How to use\n\nCLI\nAfter installation, llms_txt2ctx is available in your terminal.\nTo get help for the CLI:\nllms_txt2ctx -h\nTo convert an llms.txt file to XML context and save to llms.md:\nllms_txt2ctx llms.txt > llms.md\nPass --optional True to add the ‘optional’ section of the input file.\n\n\nPython module\n\nfrom llms_txt import *\n\n\nsamp = Path('llms-sample.txt').read_text()\n\nUse parse_llms_file to create a data structure with the sections of an llms.txt file (you can also add optional=True if needed):\n\nparsed = parse_llms_file(samp)\nlist(parsed)\n\n['title', 'summary', 'info', 'sections']\n\n\n\nparsed.title,parsed.summary\n\n('FastHTML',\n 'FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore\\'s `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.')\n\n\n\nlist(parsed.sections)\n\n['Docs', 'Examples', 'Optional']\n\n\n\nparsed.sections.Optional[0]\n\n{ 'desc': 'A subset of the Starlette documentation useful for FastHTML '\n 'development.',\n 'title': 'Starlette full documentation',\n 'url': 'https://gist.githubusercontent.com/jph00/809e4a4808d4510be0e3dc9565e9cbd3/raw/9b717589ca44cedc8aaf00b2b8cacef922964c0f/starlette-sml.md'}\n\n\nUse create_ctx to create an LLM context file with XML sections, suitable for systems such as Claude (this is what the CLI calls behind the scenes).\n\nctx = create_ctx(samp)\n\n\nprint(ctx[:300])\n\n<project title=\"FastHTML\" summary='FastHTML is a python library which brings together Starlette, Uvicorn, HTMX, and fastcore's `FT` \"FastTags\" into a library for creating server-rendered hypermedia applications.'>\nRemember:\n\n- Use `serve()` for running uvicorn (`if __name__ == \"__main__\"` is not\n\n\n\n\nImplementation and tests\nTo show how simple it is to parse llms.txt files, here’s a complete parser in <20 lines of code with no dependencies:\nfrom pathlib import Path\nimport re,itertools\n\ndef chunked(it, chunk_sz):\n it = iter(it)\n return iter(lambda: list(itertools.islice(it, chunk_sz)), [])\n\ndef parse_llms_txt(txt):\n \"Parse llms.txt file contents in `txt` to a `dict`\"\n def _p(links):\n link_pat = '-\\s*\\[(?P<title>[^\\]]+)\\]\\((?P<url>[^\\)]+)\\)(?::\\s*(?P<desc>.*))?'\n return [re.search(link_pat, l).groupdict()\n for l in re.split(r'\\n+', links.strip()) if l.strip()]\n\n start,*rest = re.split(fr'^##\\s*(.*?$)', txt, flags=re.MULTILINE)\n sects = {k: _p(v) for k,v in dict(chunked(rest, 2)).items()}\n pat = '^#\\s*(?P<title>.+?$)\\n+(?:^>\\s*(?P<summary>.+?$)$)?\\n+(?P<info>.*)'\n d = re.search(pat, start.strip(), (re.MULTILINE|re.DOTALL)).groupdict()\n d['sections'] = sects\n return d\nWe have provided a test suite in tests/test-parse.py and confirmed that this implementation passes all tests.",
"crumbs": [
"Code",
"Python module & CLI"
]
},
{
"objectID": "changes.html",
"href": "changes.html",
"title": "Changes",
"section": "",
"text": "The original llms.txt proposal was published in September 2024, when the idea that language models would routinely read websites was still speculative. Since then the community has taken to the proposal far more than I expected. Thousands of sites now publish an llms.txt file, documentation platforms generate one automatically, and coding agents use them reliably. That shift is what this revision reflects.\nAdoption brought requests, and the commonest was discoverability. Given a page, how does an agent find its markdown version, or the llms.txt file that covers it, without guessing? v2 answers with standard link relations: rel=\"alternate\" type=\"text/markdown\" points to a page’s markdown version, and rel=\"describedby\" points to the llms.txt file that covers it, provided as HTML <link> elements or an HTTP Link: header.\nPractice also diverged from v1 in ways worth blessing. v1 specified one URL form for markdown versions, .md appended to the full page URL (page.html.md). Some publishing tools instead replace the extension (page.md), so v2 allows both. v1 permitted llms.txt files in subpaths without saying what that meant. v2 defines it: a file covers the pages under its path, and the most specific file applies. This is also what lets a site that only controls a path, such as a GitHub Pages project site, participate fully.\nv1 said nothing about how llms.txt should be consumed, and described the llms_txt2ctx tool for expanding a file into an LLM context. v2 instead states the expectation directly: agents view or search the llms.txt to find what they need, then follow the relevant links, which should point to LLM-friendly content. The context-expansion tooling is no longer part of the proposal, and with it goes the special meaning of the Optional section, which told those tools what to omit. Optional sections are still allowed, and remain a useful convention for secondary links, but they no longer carry mechanical semantics. Finally, the background and examples now describe how agents actually use websites, rather than predicting that they might.",
"crumbs": [
"Changes"
]
},
{
"objectID": "changes.html#v2-august-2026",
"href": "changes.html#v2-august-2026",
"title": "Changes",
"section": "",
"text": "The original llms.txt proposal was published in September 2024, when the idea that language models would routinely read websites was still speculative. Since then the community has taken to the proposal far more than I expected. Thousands of sites now publish an llms.txt file, documentation platforms generate one automatically, and coding agents use them reliably. That shift is what this revision reflects.\nAdoption brought requests, and the commonest was discoverability. Given a page, how does an agent find its markdown version, or the llms.txt file that covers it, without guessing? v2 answers with standard link relations: rel=\"alternate\" type=\"text/markdown\" points to a page’s markdown version, and rel=\"describedby\" points to the llms.txt file that covers it, provided as HTML <link> elements or an HTTP Link: header.\nPractice also diverged from v1 in ways worth blessing. v1 specified one URL form for markdown versions, .md appended to the full page URL (page.html.md). Some publishing tools instead replace the extension (page.md), so v2 allows both. v1 permitted llms.txt files in subpaths without saying what that meant. v2 defines it: a file covers the pages under its path, and the most specific file applies. This is also what lets a site that only controls a path, such as a GitHub Pages project site, participate fully.\nv1 said nothing about how llms.txt should be consumed, and described the llms_txt2ctx tool for expanding a file into an LLM context. v2 instead states the expectation directly: agents view or search the llms.txt to find what they need, then follow the relevant links, which should point to LLM-friendly content. The context-expansion tooling is no longer part of the proposal, and with it goes the special meaning of the Optional section, which told those tools what to omit. Optional sections are still allowed, and remain a useful convention for secondary links, but they no longer carry mechanical semantics. Finally, the background and examples now describe how agents actually use websites, rather than predicting that they might.",
"crumbs": [
"Changes"
]
}
]