DarkMatter AIDC icon DARKMATTER SOFTWARE AIDC · AI Data Centers Project
Version 0.4.2 · Intelligence Platform Active

AIDC AI Data Centers Project

DarkMatter AIDC is an intelligence and knowledge-mapping platform built to investigate the global infrastructure supporting artificial intelligence. It records the organizations, facilities, services, technologies, ownership structures, energy dependencies, cloud relationships, and operational links that connect the modern AI ecosystem.

Project Overview

Mapping the invisible architecture behind artificial intelligence

AI services appear simple at the surface: a prompt enters, a response returns. Beneath that interaction is a dense industrial network of data centers, accelerator clusters, cloud platforms, power systems, cooling infrastructure, fiber routes, software stacks, corporate ownership, contractors, and financial relationships. AIDC exists to make that hidden structure visible.

What AIDC Is

AIDC is a structured intelligence environment for collecting, classifying, connecting, and reviewing information about the AI data center ecosystem. It combines a human-operated research interface with machine-assisted collection, entity modeling, relationship mapping, evidence storage, and controlled publication.

Why It Exists

AI infrastructure is distributed across companies, facilities, regions, technologies, utilities, and service providers. Public information is fragmented across corporate websites, technical documents, filings, reports, announcements, and infrastructure records. AIDC consolidates those fragments into a navigable intelligence model.

Primary Subject

The project focuses on the physical and organizational systems that make large-scale AI possible: hyperscale data centers, GPU and accelerator deployments, cloud regions, networking, energy, cooling, ownership, hosting arrangements, and supporting technologies.

Operating Principle

AIDC distinguishes between discovered information, proposed intelligence, and accepted knowledge. Research can be accelerated by automation, but permanent records remain governed by evidence, context, and human review.

Mission

Build a durable intelligence map of the AI industrial ecosystem

The project is designed to answer not only where AI data centers exist, but who owns them, who operates them, what technologies they use, what services they provide, what infrastructure they depend on, and how they connect to the wider AI economy.

Identify

Discover organizations, facilities, services, technologies, projects, regions, and infrastructure components relevant to AI data center operations.

Classify

Assign consistent types, categories, roles, technologies, domains, and operational characteristics so information can be compared across the ecosystem.

Connect

Record ownership, hosting, supply, dependency, investment, technology, energy, network, service, and geographic relationships between entities.

Verify

Preserve source URLs, excerpts, confidence, timestamps, and review decisions so claims remain traceable to evidence rather than becoming detached assertions.

Observe

Track changes in infrastructure, ownership, services, technology adoption, expansion, project status, and public reporting over time.

Export

Prepare structured information for human analysis, machine-readable outputs, future knowledge graphs, research tools, and connected DarkMatter systems.

Knowledge Architecture

A layered model built for entities, evidence, and relationships

AIDC organizes intelligence as a connected structure rather than a flat list. Every layer adds context, from broad classifications down to individual claims and the sources that support them.

01

Categories

High-level divisions such as organizations, infrastructure, technologies, services, energy, networking, ownership, regions, and projects.

02

Data Centers and Infrastructure Records

Physical facilities, campuses, cloud regions, operational sites, announced developments, compute deployments, and infrastructure locations.

03

Entities

Companies, brands, operators, landlords, utilities, cloud providers, AI laboratories, vendors, investors, technologies, and services.

04

Relationships

Ownership, operation, hosting, supply, dependency, partnership, investment, acquisition, technology use, network service, and geographic association.

05

Evidence and Review

Source documents, source URLs, excerpts, extracted claims, confidence scores, proposed changes, operator decisions, and activity history.

Intelligence Domains

The ecosystem is larger than the facility

AIDC studies each data center as part of a broader network of physical, technical, commercial, and political dependencies.

Organizations and Ownership

  • Operators, landlords, tenants, and cloud providers
  • Parent companies, subsidiaries, and acquisitions
  • Joint ventures, investors, and financing relationships
  • AI laboratories, model providers, and infrastructure partners
  • Contractors, vendors, and managed service providers

Facilities and Geography

  • Data center campuses and individual facilities
  • Cloud regions and availability zones
  • Operational, planned, proposed, and retired sites
  • Regional clustering and geographic concentration
  • Jurisdictional and infrastructure context

Compute and Technology

  • GPU, accelerator, and specialized compute platforms
  • Server, storage, orchestration, and networking technologies
  • Cloud services, hosting platforms, and AI infrastructure stacks
  • Cooling systems and facility management technologies
  • Security, monitoring, automation, and operational software

Energy and Physical Dependencies

  • Grid connections, utilities, generation, and substations
  • Power purchase agreements and renewable energy claims
  • Cooling, water, and environmental dependencies
  • Fiber, transit, interconnection, and network routes
  • Land, construction, logistics, and regional capacity constraints
Operational Model

From discovery to accepted intelligence

AIDC combines human investigation, structured interfaces, and controlled automation. Each stage has a defined role so that collection never silently becomes publication.

01 / DISCOVER

Find

Identify organizations, facilities, technologies, services, documents, projects, and candidate relationships.

02 / COLLECT

Retrieve

Gather source material from official pages, public records, documentation, reports, filings, and operator-provided research targets.

03 / STRUCTURE

Model

Convert raw information into entities, attributes, classifications, relationships, evidence records, and proposed changes.

04 / REVIEW

Evaluate

Inspect source quality, context, confidence, duplication, relevance, and compatibility with the existing intelligence model.

05 / PUBLISH

Accept

Commit approved information to the AIDC knowledge base while preserving evidence, review state, and operational history.

06 / EXPLORE

Analyze

Navigate the resulting hierarchy and relationship network to understand the architecture behind AI infrastructure.

Platform Capabilities

A research console, knowledge base, and operational map

Browse Navigate categories, records, and entities
Connect Map ownership, services, and dependencies
Review Control what becomes accepted intelligence
Expand Grow the model through research and automation
Integrated Research Module

DarkMatter AIDC Harvester

Harvester is the intelligence gathering crawler built for AIDC. It turns research targets into controlled jobs, discovers and retrieves sources, extracts structured claims, scores evidence, stages proposed database changes, and routes every publication decision through a human review gate.

DarkMatter AIDC Harvester icon
Harvester · AIDC 0.4.2

Automated Intelligence Collection Engine

Harvester extends AIDC without bypassing its governance model. It accelerates repetitive collection and extraction work while preserving source provenance, confidence, activity logs, proposal status, and operator authority.

0.4.2 Processing Architecture

From controlled collection to publishable intelligence

The current project separates execution, extraction, review, and publication into distinct responsibilities. This keeps crawler success from being confused with verified knowledge and prevents endpoint data from leaking into generic relationship records.

worker.php · orchestration

Claims queued research jobs, enforces process locking and runtime limits, retrieves source documents, records execution events, invokes extraction, and leaves every proposed database change inside the review boundary.

extractor.php · structured interpretation

Sends bounded document context and the research goal to Ollama, requires schema-shaped JSON, validates canonical names, records the extraction engine and source endpoint, and rejects the former regex behavior that could turn sentence fragments into entities.

Review-safe publication

Publishing description, website, sector, company type, or a primary company proposal can resolve or create the canonical company first. Review order therefore does not orphan enrichment data, while duplicate-safe resolution protects existing records.

Correct relationship placement

Endpoints publish to endpoints; direct company links publish to company_endpoints; technologies and services use their dedicated relation layers; ownership and other graph assertions remain in the relationship model.

Governance

Evidence before authority

AIDC is built around a deliberate distinction between collected material and accepted knowledge. The system can assist discovery, correlation, and proposal generation, but publication remains a controlled decision.

Human-controlled publication

Automated tools may gather evidence and stage candidate changes, but they do not receive unrestricted authority to overwrite the primary intelligence model.

Source traceability

Intelligence should remain connected to its originating source, excerpt, retrieval record, confidence, and review history wherever possible.

Structured uncertainty

Incomplete, conflicting, or uncertain information can be retained as research material without being falsely promoted to verified fact.

Operational history

Changes, worker activity, review actions, errors, and publication decisions create a record that supports debugging, auditing, and future reassessment.

Conceptual Model

The AIDC intelligence chain

Artificial Intelligence Service
    │
    ├── Model Provider
    │     ├── Organization
    │     ├── Cloud Platform
    │     └── Compute Dependency
    │
    ├── Data Center / Cloud Region
    │     ├── Owner
    │     ├── Operator
    │     ├── Tenant
    │     ├── Location
    │     └── Project Status
    │
    ├── Infrastructure
    │     ├── GPU / Accelerator Platform
    │     ├── Network / Interconnection
    │     ├── Energy / Utility
    │     ├── Cooling / Water
    │     └── Security / Operations
    │
    └── Evidence
          ├── Source Document
          ├── Extracted Claim
          ├── Confidence
          ├── Review Decision
          └── Publication Record
Version 0.4.2 Documentation

Project records for operators, developers, and machines

The 0.4.2 documentation set joins the original AIDC ownership-map canon with the current Harvester implementation, corrected endpoint publication model, provenance controls, and machine-readable data strategy.

Machine-readable surfaces

AIDC exposes structured JSON API routes and database-backed entity records so companies, endpoints, technologies, services, brands, evidence, and relationships can be consumed without scraping presentation HTML. Human pages remain the visual intelligence console; machine outputs remain explicit data contracts.

  • JSON endpoint and company records
  • Canonical company-to-endpoint links
  • Source URLs and extraction provenance
  • Structured claims and review states