Web data for AI training and generative search

Web data with a chain of evidence.

Every page we fetch comes back with a receipt: where it came from, what was checked, which permission signals were present, and exactly what it cost. Pay as you go, from $0.00012 per page.

$0.00012
Basic crawl
$0.00050
Standard SERP
$0
Minimum top-up
SAMPLE RECEIPT / JSON
{
  "receipt_id": "rcpt_8f92c1_001",
  "source_url": "https://developer.mozilla.org/en-US/docs/Web",
  "fetched_at": "2026-09-27T10:42:18Z",
  "fetch_method": "http_html",
  "http_status": 200,
  "content_type": "text/html",
  "content_sha256": "b92c4f…7a05d3",
  "content_check": "substantive",
  "robots_txt": "allowed",
  "ai_use_signal": "review_required",
  "billed_pages": 1,
  "unit_cost_usd": 0.00012
}

Sample receipt · returned with every crawled page

Product overview · 14 seconds

Try it

Fetch a real page now.

Paste any public URL to see the markdown our pipeline extracts. Sign up to get the full receipt, robots.txt evidence, and API access.

Free, rate-limited, and not stored. Returns the first 3,000 characters of extracted markdown.

The problem

Most web data arrives with no idea of where it came from.

Unverifiable provenance

A row in a dataset rarely says which URL it came from, when it was fetched, or whether the page has changed since.

Unclear permissions

Robots directives and AI-use declarations are checked once, if at all, and almost never stored with the data itself.

Opaque cost

Credit bundles, minimum top-ups, and per-task pricing hide what a single page actually costs you.

The solution

We attach the evidence to the data, so an audit is a lookup rather than an investigation.

Provenance

URL, timestamp, method, status, and a content hash you can re-verify.

Permission signals

What robots.txt and AI-use declarations said at the moment of the fetch.

Cost per unit

The billable units and exact charge for that page, in the same record.

How it works

Six stages, all of them recorded.

Read the receipt specification
  1. 01

    Resolve

    Normalize the target URL, resolve the host, and record the exact request that will be made.

  2. 02

    Fetch

    Retrieve the page with a declared user agent, capturing status, headers, and timing.

  3. 03

    Normalize

    Extract substantive content and drop boilerplate, then hash the result for verification.

  4. 04

    Check permissions

    Read robots.txt and machine-readable AI-use signals for the agent and path.

  5. 05

    Receipt

    Write provenance, checks, signals, and cost into one inspectable record.

  6. 06

    Deliver

    Return content and receipt together, in JSON your pipeline can store and audit.

Coverage

What you can pull, and how fast.

Page crawl

Available

Content, extraction context, and a receipt for a single URL or a batch of up to 100.

Search results

Available

Organic results with position, domain, and the same receipt attached to each fetch.

Site audit crawl

In build

Crawl a site you own and get per-page issues alongside the evidence trail.

AI answer visibility

On the roadmap

Track whether assistants cite a brand, and which sources they name.

Live

One request, answer inline, for anything a person is waiting on.

Batch

Up to 100 targets per request, so volume stays cheap.

Scheduled

Recurring jobs delivered to your webhook, no polling.

Published rate card

Published unit prices, no minimum top-up.

Rates below are our published pay-as-you-go proposal. You pay per successful unit, and every unit comes with its receipt.

UnitRatePer 1,000
Basic crawled page$0.00012$0.12
Standard SERP (10 results)$0.00050$0.50
Minimum top-up$0—

Rates are a published proposal pending commercial approval. Failed and blocked fetches are never billed.

ISO 27001 certified

Guni Innovations Pte. Ltd. is certified for information security management. Request the certificate or a DPA at any time.

Singapore incorporated

UEN 202302437E, operating under the laws of the Republic of Singapore, with a named contact for legal and security review.

No invented proof

We publish only claims we can evidence. You will not find borrowed logos or unverifiable scale numbers on this site.

Answers

Frequently asked questions

Still unsure? Write to dani@dataforgaio.com or read the full FAQ.

Put a receipt behind every page you use.

Start free with no deposit, explore the dashboard, or ask us about volume rates and procurement.