Company Tech DNA — Multi-Source Technology Profiles, AI Maturity & Peer Similarity
3,171 deep company technology profiles fused from four independent evidence layers: JD skill tags from 394,300 job postings, 495,363-domain website technology fingerprints, LLM-labeled hiring adoption signals, and US patent CPC aggregates. Cloud selection, AI maturity (0–100), tech-debt heuristics, research directions and top-10 peer similarity — every score reproducible with a shipped methodology.
- Records
- 124,826 records
- Refresh
- Monthly refresh (subscription)
- Formats
- CSV + Parquet
- Coverage
- Global (US-anchored, tech/startup skew)
- Packages
- 3
- Total volume
- 6.1 MB
Slices from
$199
Request access / Get a sampleSelf-serve checkout is coming online — request access and we reply within one business day.
What you get
- S — AI-track pack: all 327 companies with AI maturity ≥ 40, profiles + evidence matrix + peers
- M — All 3,171 company profiles (cloud mix, AI maturity, tech debt, patent research directions) + top-10 peer similarity
- L — Everything in M plus the full 93,965-row company × technology × source evidence matrix
- METHODOLOGY.md in every package: exact formulas and source row counts for every score
- 2,517 of 3,171 companies observed by 2+ independent sources; 877 with patent DNA
Sample preview
First 50 rows of the real free-sample file — the identical schema you get in every package. Click a column to sort, or filter across all rows.
| 1.3k | uncountable | 6 | 2 | 6 | 10.0 | ||||||
| 2k.com | 2K | 56 | 3 | 186 | 122.0 | aws | 0.579 | 0.263 | 0.158 | 34.8 | |
| 360learning.com | 360learning | 23 | 2 | 52 | 30.0 | azure | 0.0 | 0.0 | 1.0 | 17.7 | |
| 9am.health | join9am | 39 | 2 | 106 | 18.0 | aws | 1.0 | 0.0 | 0.0 | 26.0 | |
| a16zcrypto.com | a16z-crypto | 9 | 1 | 9 | 1.0 | 0.0 | |||||
| aarons.com | Dev | Staffing And Recruiting | 14 | 2 | 17 | azure | 0.0 | 0.0 | 1.0 | 0.0 | |
| abar.tech | infracost | 6 | 2 | 7 | 1.0 | azure | 0.0 | 0.0 | 1.0 | 19.3 | |
| abridge.com | abridge | 50 | 3 | 273 | 42.0 | gcp | 0.318 | 0.636 | 0.045 | 54.3 | |
| absorblms.com | absorblms | 56 | 3 | 143 | 33.0 | aws | 0.5 | 0.333 | 0.167 | 28.4 | |
| abusix.com | abusix | 6 | 2 | 12 | 2.0 | 30.0 |
50 rows · page 1 of 5
Download the free sample pack
Real rows with the identical schema — enter your work email and we issue a download link instantly. No credit card, no spam.
Data dictionary summary
Key fields — the full data dictionary ships with every purchase.
| field | type | description |
|---|---|---|
| company_domain | string | Company website domain (join key across all sources) |
| ai_maturity | float | 0–100 composite: JD AI-skill share + LLM-labeled adoption + AI team building + AI API usage + AI patents |
| primary_cloud | string | aws / gcp / azure with the most hiring evidence, plus cloud_mix_* shares |
| tech_debt | float | Share (%) of the company's techs in a documented legacy list (with legacy_techs naming them) |
| top_cpc_subclasses | string | Top-5 patent CPC subclasses = in-research directions (877 companies) |
| differentiation | float | 100 × (1 − mean top-5 peer similarity): how unlike its closest peers the stack is |
| tech_matrix.evidence_count | int | Per company × technology × source observations, with a detail field naming the basis |
| peer_similarity.similarity | float | TF-IDF-weighted cosine similarity, top-10 peers per company |
Coverage statistics
- Records
- 124,826
- Company profiles
- 3,171
- Fields
- 33
- Time coverage
- n/a
- Largest package
- 3.6 MB
Available packages
Real contents of the current builds — row counts, sizes and checksums come from the packaging manifest, not estimates.
| Slice | Scope | Rows | Size | Snapshot | Monthly |
|---|---|---|---|---|---|
| S — ai-track-profiles | Buyer-scenario pack: every company with AI maturity ≥ 40 (327 companies) — profiles + full evidence matrix + peer similarity, for AI-infra vendors and investors scanning the AI-adopter track | 23,499 | 638 KB | $199 | $99/mo |
| M — profiles-and-similarity | All 3,171 company profiles (cloud mix, AI maturity, tech debt, research directions) + top-10 peer similarity table — the scores layer without the row-level evidence matrix | 30,861 | 1.9 MB | $499 | $199/mo |
| L — full-matrix | Everything in M plus the full company × technology × source evidence matrix (93,965 rows) — rebuild every score yourself or run your own similarity models | 124,826 | 3.6 MB | $999 | $399/mo |
Zero new collection: cross-analysis of four archived DataForge lines — 394,300 ATS job postings (JD skill tags), 495,363-domain website technology fingerprints, 115,976 company hiring signals, and per-company US patent CPC aggregates — resolved to 3,171 deep company profiles (2,517 companies with 2+ independent sources, 877 with patent DNA).
Full package inventory
Every package registered in the delivery catalog for this dataset — 3 packages, 6.1 MB total, 179,186 rows. Sizes, row counts and SHA-256 checksums come straight from the packaging pipeline.
| Package | Date | Rows | Size | sha256 | Access |
|---|---|---|---|---|---|
| Loading 3 packages… | |||||
How to get this data
DataForge (this site)
External channels
No free external mirror yet — this dataset is delivered exclusively through DataForge. See the Open Data program →
# Machine-readable package index (size / sha256 / rows / channels)
curl -s https://data.zalize.com/datasets/company-tech-dna-profiles-dataset/packages.json | jq '.packages[0]'
# Request a free sample (signed link sent to your email)
curl -X POST https://dl.zalize.com/api/sample \
-H 'Content-Type: application/json' \
-d '{"dataset":"company-tech-dna-profiles-dataset","email":"you@company.com"}'Pricing
One-time snapshot or monthly subscription. Prices are for the Internal Analytics tier; AI-training and redistribution tiers available — see License. Tiers marked “Request access” require a short approval step before delivery.
| Tier | Scope | One-time snapshot | Subscription | Access |
|---|---|---|---|---|
| S — ai-track-profiles | Buyer-scenario pack: every company with AI maturity ≥ 40 (327 companies) — profiles + full evidence matrix + peer similarity, for AI-infra vendors and investors scanning the AI-adopter track | $199 | $99/mo | Self-serve |
| M — profiles-and-similarityPOPULAR | All 3,171 company profiles (cloud mix, AI maturity, tech debt, research directions) + top-10 peer similarity table — the scores layer without the row-level evidence matrix | $499 | $199/mo | Self-serve |
| L — full-matrix | Everything in M plus the full company × technology × source evidence matrix (93,965 rows) — rebuild every score yourself or run your own similarity models | $999 | $399/mo | Self-serve |
AI-training use requires the AI-Training tier (1.5× base); redistribution licensing is quoted individually. Annual subscriptions get 2 months free. License details →
Update frequency & delivery
- Zero new collection: pure cross-analysis of four archived DataForge lines (jobs, techstack, hiring signals, patents), rebuilt monthly.
- Patent matching is deliberately conservative (audited 0.47% mis-merge method): absence means no confident match, not no patents.
- No JD text, no patent text, no personal data — company-level derived observations only.
Compliance & provenance
Collected only from publicly accessible pages and endpoints — no login-required data, no CAPTCHA solving, no forged authentication. Collection is rate-limited with an identified User-Agent, and collection logs are retained. PII is removed or pseudonymized. Every record carries its source URL and collection timestamp. Sold under a tiered commercial license.
FAQ
How is this different from BuiltWith or Wappalyzer?
Website fingerprints only see client-visible technologies. Tech DNA fuses fingerprints with hiring evidence (JD skills + LLM-labeled adoption signals) and patent classifications, so it covers backend stacks, cloud selection, AI maturity and research directions — with per-source evidence counts. Enterprise alternatives (HG Insights) start at $3k–12k+/yr; BuiltWith at $2,950/yr.
How are the scores computed?
Every package ships METHODOLOGY.md with the exact formula and the source table + row counts behind each component. The L tier includes the full evidence matrix so you can rebuild every score yourself.
How is the data delivered?
Self-serve packages are delivered instantly after checkout via a time-limited download link from our delivery service (dl.zalize.com).
Can I use this data for AI training?
Yes — with the AI-Training license tier (1.5× the base price). See the License page for details.
Related datasets
Hiring Signals
115,976 evidence-backed hiring signals across 6,295 companies, derived from 394,300 live ATS job postings: which companies are adopting which technologies, building AI teams, expanding into new cities or hiring first leadership roles. Every signal carries up to 5 evidence posting URLs. No job-description text included.
- Records
- —
- Refresh
- Monthly
- From
- $99
Tech Stack by Domain
Detected technology stacks for 495,363 websites — frameworks, analytics, CMS, hosting and e-commerce platforms per domain — with a change-tracking snapshot pair for migration detection. Available to B2B buyers on request.
- Records
- 495,363
- Refresh
- On
- From
- Contact
App Category Opportunity Scan
A ranked league table of 64 app-store categories scored by demand vs. user dissatisfaction — where download volume is high but ratings are weak — with 348 representative high-demand low-rated apps as concrete build targets, for both Apple App Store and Google Play.
- Records
- 412
- Refresh
- Monthly
- From
- $99