Skip to content
DataForge
Sales IntelligenceFresh — updated 2026-08-06

Company Tech DNA — Multi-Source Technology Profiles, AI Maturity & Peer Similarity

3,171 deep company technology profiles fused from four independent evidence layers: JD skill tags from 394,300 job postings, 495,363-domain website technology fingerprints, LLM-labeled hiring adoption signals, and US patent CPC aggregates. Cloud selection, AI maturity (0–100), tech-debt heuristics, research directions and top-10 peer similarity — every score reproducible with a shipped methodology.

Records
124,826 records
Refresh
Monthly refresh (subscription)
Formats
CSV + Parquet
Coverage
Global (US-anchored, tech/startup skew)
Packages
3
Total volume
6.1 MB

Slices from

$199

Request access / Get a sample

Self-serve checkout is coming online — request access and we reply within one business day.

What you get

  • S — AI-track pack: all 327 companies with AI maturity ≥ 40, profiles + evidence matrix + peers
  • M — All 3,171 company profiles (cloud mix, AI maturity, tech debt, patent research directions) + top-10 peer similarity
  • L — Everything in M plus the full 93,965-row company × technology × source evidence matrix
  • METHODOLOGY.md in every package: exact formulas and source row counts for every score
  • 2,517 of 3,171 companies observed by 2+ independent sources; 877 with patent DNA

Sample preview

First 50 rows of the real free-sample file — the identical schema you get in every package. Click a column to sort, or filter across all rows.

Sample rows from Company Tech DNA
1.3kuncountable62610.0
2k.com2K563186122.0aws0.5790.2630.15834.8
360learning.com360learning2325230.0azure0.00.01.017.7
9am.healthjoin9am39210618.0aws1.00.00.026.0
a16zcrypto.coma16z-crypto9191.00.0
aarons.comDevStaffing And Recruiting14217azure0.00.01.00.0
abar.techinfracost6271.0azure0.00.01.019.3
abridge.comabridge50327342.0gcp0.3180.6360.04554.3
absorblms.comabsorblms56314333.0aws0.50.3330.16728.4
abusix.comabusix62122.030.0

50 rows · page 1 of 5

Download the free sample pack

Real rows with the identical schema — enter your work email and we issue a download link instantly. No credit card, no spam.

Data dictionary summary

Key fields — the full data dictionary ships with every purchase.

Field definitions for Company Tech DNA
fieldtypedescription
company_domainstringCompany website domain (join key across all sources)
ai_maturityfloat0–100 composite: JD AI-skill share + LLM-labeled adoption + AI team building + AI API usage + AI patents
primary_cloudstringaws / gcp / azure with the most hiring evidence, plus cloud_mix_* shares
tech_debtfloatShare (%) of the company's techs in a documented legacy list (with legacy_techs naming them)
top_cpc_subclassesstringTop-5 patent CPC subclasses = in-research directions (877 companies)
differentiationfloat100 × (1 − mean top-5 peer similarity): how unlike its closest peers the stack is
tech_matrix.evidence_countintPer company × technology × source observations, with a detail field naming the basis
peer_similarity.similarityfloatTF-IDF-weighted cosine similarity, top-10 peers per company

Coverage statistics

Records
124,826
Company profiles
3,171
Fields
33
Time coverage
n/a
Largest package
3.6 MB

Available packages

Real contents of the current builds — row counts, sizes and checksums come from the packaging manifest, not estimates.

Built packages for Company Tech DNA
SliceScopeRowsSizeSnapshotMonthly
S — ai-track-profilesBuyer-scenario pack: every company with AI maturity ≥ 40 (327 companies) — profiles + full evidence matrix + peer similarity, for AI-infra vendors and investors scanning the AI-adopter track23,499638 KB$199$99/mo
M — profiles-and-similarityAll 3,171 company profiles (cloud mix, AI maturity, tech debt, research directions) + top-10 peer similarity table — the scores layer without the row-level evidence matrix30,8611.9 MB$499$199/mo
L — full-matrixEverything in M plus the full company × technology × source evidence matrix (93,965 rows) — rebuild every score yourself or run your own similarity models124,8263.6 MB$999$399/mo

Zero new collection: cross-analysis of four archived DataForge lines — 394,300 ATS job postings (JD skill tags), 495,363-domain website technology fingerprints, 115,976 company hiring signals, and per-company US patent CPC aggregates — resolved to 3,171 deep company profiles (2,517 companies with 2+ independent sources, 877 with patent DNA).

Full package inventory

Every package registered in the delivery catalog for this dataset — 3 packages, 6.1 MB total, 179,186 rows. Sizes, row counts and SHA-256 checksums come straight from the packaging pipeline.

PackageDateRowsSizesha256Access
Loading 3 packages…

How to get this data

External channels

No free external mirror yet — this dataset is delivered exclusively through DataForge. See the Open Data program →

# Machine-readable package index (size / sha256 / rows / channels)
curl -s https://data.zalize.com/datasets/company-tech-dna-profiles-dataset/packages.json | jq '.packages[0]'

# Request a free sample (signed link sent to your email)
curl -X POST https://dl.zalize.com/api/sample \
  -H 'Content-Type: application/json' \
  -d '{"dataset":"company-tech-dna-profiles-dataset","email":"you@company.com"}'

Pricing

One-time snapshot or monthly subscription. Prices are for the Internal Analytics tier; AI-training and redistribution tiers available — see License. Tiers marked “Request access” require a short approval step before delivery.

Pricing for Company Tech DNA
TierScopeOne-time snapshotSubscriptionAccess
S — ai-track-profilesBuyer-scenario pack: every company with AI maturity ≥ 40 (327 companies) — profiles + full evidence matrix + peer similarity, for AI-infra vendors and investors scanning the AI-adopter track$199$99/moSelf-serve
M — profiles-and-similarityPOPULARAll 3,171 company profiles (cloud mix, AI maturity, tech debt, research directions) + top-10 peer similarity table — the scores layer without the row-level evidence matrix$499$199/moSelf-serve
L — full-matrixEverything in M plus the full company × technology × source evidence matrix (93,965 rows) — rebuild every score yourself or run your own similarity models$999$399/moSelf-serve

AI-training use requires the AI-Training tier (1.5× base); redistribution licensing is quoted individually. Annual subscriptions get 2 months free. License details →

Update frequency & delivery

  • Zero new collection: pure cross-analysis of four archived DataForge lines (jobs, techstack, hiring signals, patents), rebuilt monthly.
  • Patent matching is deliberately conservative (audited 0.47% mis-merge method): absence means no confident match, not no patents.
  • No JD text, no patent text, no personal data — company-level derived observations only.

Compliance & provenance

Collected only from publicly accessible pages and endpoints — no login-required data, no CAPTCHA solving, no forged authentication. Collection is rate-limited with an identified User-Agent, and collection logs are retained. PII is removed or pseudonymized. Every record carries its source URL and collection timestamp. Sold under a tiered commercial license.

FAQ

How is this different from BuiltWith or Wappalyzer?

Website fingerprints only see client-visible technologies. Tech DNA fuses fingerprints with hiring evidence (JD skills + LLM-labeled adoption signals) and patent classifications, so it covers backend stacks, cloud selection, AI maturity and research directions — with per-source evidence counts. Enterprise alternatives (HG Insights) start at $3k–12k+/yr; BuiltWith at $2,950/yr.

How are the scores computed?

Every package ships METHODOLOGY.md with the exact formula and the source table + row counts behind each component. The L tier includes the full evidence matrix so you can rebuild every score yourself.

How is the data delivered?

Self-serve packages are delivered instantly after checkout via a time-limited download link from our delivery service (dl.zalize.com).

Can I use this data for AI training?

Yes — with the AI-Training license tier (1.5× the base price). See the License page for details.

Slices from

$199

Request access