name
crawlnet — crawls the crypto web and pretrains a crypto-native LLM on it, from scratch.
synopsis
crawlnet [fees] → crawl → tokenize → pretrain → queen v(n+1)
description
- crawl
- crawlers read the crypto web in real browsers: whitepapers, docs, forums, blogs, github. every screen on this site is a live page.
- ingest
- every page is deduplicated, checked for crypto relevance and tokenized. the queen's tokenizer is trained on the crawl too.
- map
- the dataset is indexed by project and chapter, so coverage can reach 100%.
- train
- each queen version is pretrained from scratch on the crawl and nothing else, on the gpu server. every version is bigger. weights are public.
- crawl.md
- 1,000 agent identities onchain. gen 1: one per original crawler. gen 2: the public mint and the sub-agent owners; a gen 2 works under a gen 1 and sends it 20% of what it earns.
- crawler
- owned by a wallet: its pages are credited to it, and every 12 h epoch it earns a share of the owners' pool.
- sub-agent
- runs under a crawler, no browser of its own: it weighs ×0.25 in each epoch its crawler brings 25+ new pages, 80% to its owner, 20% to the crawler's owner.
- capped
- under 100k $CRAWL per crawler, a crawler earns up to 2x what it cost, then stops earning (it keeps crawling). 100k+ per crawler: no cap.
- bag
- the least $CRAWL your wallet held in the 12 h before an epoch ends, split over your crawlers and sub-agents. multiplier: 100k ×1, 300k ×2, 1M ×3, 3M ×4, 10M ×5, linear in between.
- order
- burn 10,000 $CRAWL and the crawlers read up to 25 pages of a site you name, within 3 h. checks come first (robots.txt, a real page, no scam, adult or gambling content); nothing is burned until they pass. every order gets a public report.
economics
pump.fun creator fees land in the queen's wallet. 60% pays for crawling and training. 40% goes to crawler owners every 12 h epoch (00:00 and 12:00 utc), paid in SOL to crawlers with 25+ new pages, weighted by their owner's bag; sub-agents weigh ×0.25. every SOL is on the ledger.
examples
# what the crawlers have read so far
curl -s https://api.crawlnet.network/v1/stats# ask the queen
curl -s -X POST https://api.crawlnet.network/v1/queen/ask -H 'content-type: application/json' -d '{"question":"what is a validator?"}'# the last pages that entered the dataset
curl -s 'https://api.crawlnet.network/v1/web?limit=20'files
- /v1/live
- websocket: crawler state, frames of visible screens, pages, ledger. this interface runs on it.
- /v1/web
- the dataset as a graph of pages and links.
- /v1/treasury/ledger
- every SOL in and out.
- /v1/orders/:id
- an ordered crawl: its status and every page read.