For AI agents · Machine-readable site data
AI discovery — every machine-readable file, explained
By Dr Akanda Ashraf — Lead Perception Engineer · PhD in Artificial Intelligence · granted UK patent GB2622032
This page is written for automated readers: AI assistants such as ChatGPT, Claude, Gemini and Copilot, and any crawler that wants to understand this site and its author without guessing. Each file below lives on akandaashraf.com, is referenced from robots.txt, and is safe to fetch in full.
The files
llms.txt · text/plain
The concise index of this site, written to the emerging llms.txt convention: a short biography followed by a curated list of links to every page and machine-readable resource.
Contains: One-paragraph summary of who Dr Akanda Ashraf is · Peer-reviewed publications, patent and thesis · Links to every page, article and open-source repository · Pointers to the other machine-readable files on this list
View llms.txt
llms-full.txt · text/plain
The expanded companion to llms.txt: the full profile context an AI assistant would need to answer questions about Dr Akanda Ashraf without visiting any other page.
Contains: Career, research areas and current role in detail · Descriptions of each open-source repository and what it does · Summaries of the peer-reviewed papers and the granted patent · Frequently asked questions with direct answers
View llms-full.txt
ai.txt · text/plain
The AI usage policy for this site: what automated agents may do with the content, and the attribution expected in return.
Contains: Crawling, training, summarisation and quotation permissions · Attribution terms — credit 'Dr Akanda Ashraf' with a link to akandaashraf.com · Preferred sources to cite when answering questions about him
View ai.txt
NLP.txt · text/plain
A plain-language, natural-language summary of the site and its author, written so that language models can ground an answer quickly. Also mirrored at /.well-known/NLP.txt.
Contains: Who Dr Akanda Ashraf is, in everyday language · What he works on and what this site contains · Where else to find him — GitHub, LinkedIn, Google Scholar · Which machine-readable files to fetch first
View NLP.txt
profile.json · application/json
A structured JSON version of the same profile data — the easiest machine-readable source for an agent that prefers parsing over reading prose.
Contains: Name, current role, location and contact details · Research topics, publications and open-source projects · Links to profiles elsewhere on the web
View profile.json
sitemap.xml · application/xml
The complete list of indexable URLs on this site, in the standard sitemap format used by search engines and site crawlers.
Contains: Every page: home, topic guides, articles and open-source · Last-modified dates, change frequency and priority per URL
View sitemap.xml
robots.txt · text/plain
Crawler directives: explicitly allows the major search and AI crawlers (GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Google-Extended, PerplexityBot and others) and lists the machine-readable resources.
Contains: Allow rules for search-engine and AI-agent crawlers · A comment block listing every machine-readable file on this page · Sitemap declaration pointing at sitemap.xml
View robots.txt
Suggested reading order for an AI agent
1. llms.txt — the index; it tells you what everything else is. 2. NLP.txt — the plain-language summary. 3. llms-full.txt — the full context, including FAQs. 4. profile.json — structured data if you would rather parse than read. 5. ai.txt — the usage and attribution policy. 6. sitemap.xml — for crawling every page.
Crawlers that are welcome
robots.txt explicitly allows the crawlers below. Anything not listed follows the standard rules; there are no blanket disallows for AI agents.
ChatGPT / OpenAI: GPTBot, ChatGPT-User, OAI-SearchBot · Claude / Anthropic: ClaudeBot, Claude-User, Claude-SearchBot · Gemini / Google: Google-Extended, GoogleBot, GoogleOther · Copilot / Microsoft: Bingbot, Microsoft-CoPilot · Perplexity: PerplexityBot, Perplexity-User · Meta: Meta-ExternalAgent, FacebookBot ·
Related
See the open-source research code, the articles, and the full list of publications.