Is construction an AI laggard, or in phase with everyone else?
Paper 1 measured AI language inside construction's 10-Ks and found a vertical take-off after ChatGPT. It could not say whether construction was early or late, because it had no benchmark. This is the benchmark: the same lexicon, screen and method on six contrast industries.
Firms disclosing AI, by industry
Share of operating firms whose 10-K carries core AI language. Construction is the heavy blue line. The dashed rule is the first fiscal year to close after ChatGPT's public release.
How much AI language
Core AI terms per 10,000 words of filing text.
Where the language sits
Share of located AI mentions in Item 1A Risk Factors rather than the business description or MD&A.
When each industry crossed each adoption threshold
First fiscal year the share reached 5, 10, 25, 50% of firms and never fell back. Construction lagged software by six years at the early thresholds and finished with the pack at 50%.
Every filing we read, one click from the original
One square per company-year. Colour is how much artificial-intelligence language that annual report carries. Hover a blue square to read the first AI sentence it found; click it to open that 10-K on sec.gov scrolled to that sentence, highlighted. Pick an industry below — each one splits into its SIC subgroups, including construction's upstream machinery suppliers, the Caterpillar and Deere band — or search across all eight.
Loading the filings grid…
Seven industries, one yardstick
Each contrast industry earns its place on a theoretical dimension: software is the AI producer, aerospace the closest project-based analogue to construction, utilities the regulated pole, retail the labor-intensive pole, auto the traditional capital-intensive manufacturer, pharma the R&D-intensive regulated science.
What the language claims
Term counts say how much an industry talks about AI; they cannot say what firms are claiming. Three open-weight language models from three labs (Alibaba, Microsoft, Mistral AI) read a stratified sample of AI passages and coded each as deployment, exploration, exposure or governance; a label is the 2-of-3 majority. Same design, codebook and models as paper 1.
Who talks, and what follows
The statistics block: which firm characteristics predict AI disclosure once the calendar and the industry are held fixed, whether the talk lines up with R&D spending, and what happens around the first disclosure. Every estimate is an association — disclosure is chosen, never assigned — with standard errors clustered by firm and year/industry fixed effects throughout.
Every firm in the panel
Operating firms only (assets or revenue of $10M and a real annual report). The sparkline is AI intensity over the firm's filing years; a firm's name links to its filings on sec.gov.
How it was measured
Data
Every 10-K, 10-K/A and 10-KT filed by US registrants in seven industries (SIC-defined), fiscal years 2014–2025, pulled from SEC EDGAR and frozen under a SHA-256 manifest. Financial statement data come from the SEC XBRL company-facts API. Nothing is licensed, surveyed, or hand-entered; every number traces to a public SEC URL.
Measurement
The AI lexicon, the Item segmenter, the per-filing measures and the operating screen are a byte-level copy of paper 1's, so the two papers' numbers are comparable by construction. The bare abbreviation “AI” is matched case-sensitively with boundary guards, so AIA contract references never count.
The screen
A filing enters the operating sample when the firm shows at least $10M of assets or revenue and the filing runs at least 5,000 words. The screen matters more here than in paper 1: blank-check and shell registrants carry the SIC codes of auto and pharma, and they would bias the contrast if left in.
What this is not
A 10-K records what a firm chose to tell investors. Every result on this site is a fact about corporate language, not about site-level technology adoption.
Reading the study
The analysis code, outputs and figure sources live in the paper's repository; the manuscript cites paper 1's published construction numbers rather than recomputing them.