Skip to content
DataForge

About DataForge

DataForge turns publicly accessible web data into clean, documented, ready-to-use datasets. We focus on the gap the incumbents leave open: enterprise data vendors start at $1,000+ per dataset and $2,000+/month subscriptions, while most developers, analysts and AI teams just need a well-documented slice.

What we collect

  • Job postings — directly from public ATS endpoints (Greenhouse, Lever, Ashby), with parsed salary ranges.
  • App store data — public App Store and Google Play metadata with derived review sentiment and topic metrics.
  • E-commerce catalogs — Shopify DTC store products via public endpoints, snapshotted weekly into price-history series.

Methodology

Automated collection from publicly accessible pages and endpoints only. We never collect login-walled or paywalled data, never solve CAPTCHAs or forge authentication, and crawl rate-limited with an identified User-Agent. Every record carries its source URL and collection timestamp, and collection logs are retained for auditability.

Quality & documentation

Every dataset ships with a data dictionary, a datasheet documenting sources, methods and known limitations, and its license. We publish honest quality metrics — fill rates, parse accuracy, freshness — rather than hiding them.

Privacy & compliance

Personal identifiers are removed or pseudonymized (emails and phone numbers masked, user handles pseudonymized). Copyrighted expressive content is sold only as derived metrics, except where explicitly licensed for AI training with clear source and usage disclosure. Buyers act as independent data controllers under GDPR; erasure-request pass-through is included.

Where to buy

Datasets are sold via Gumroad (instant checkout) with free samples on Kaggle and Hugging Face. B2B subscriptions and custom slices are available on request.