It is 6:40 on a Tuesday night in a field office. A project manager has a scope letter to write, a bid tab open in one window, and a free chatbot open in the other. He pastes the tab in, asks for a clean paragraph, and gets one. Nothing about that moment feels like a decision. It feels like saving twenty minutes.
Here is what actually moved. Six subcontractor quotes with unit prices that those subs shared under a promise. The owner's budget line from the last meeting. Two names and a phone number. All of it now sits on a server the company has no contract with, under terms the PM did not read, in a tier of the product that, by default, uses what you type to train the next version.
Privacy in the AI era is not a question about whether the machine is smart. It is a question about custody. Who holds the copy, under what promise, and for how long. Construction is unusually exposed on that question, because almost nothing a contractor holds is its own.
A contractor is a warehouse of other people's secrets
Read the contracts you already sign. AIA A201-2017, Section 2.2.4, says that where the owner designates information as confidential, "the Contractor shall keep the information confidential and shall not disclose it to any other person." The carve-out is for people who "need to know the content of such information solely and exclusively for the Project and who agree to maintain the confidentiality." A consumer chatbot has agreed to nothing.
Section 1.5.2 of the same document says the contractor and its subs may reproduce the drawings and project manual "solely and exclusively for execution of the Work." The design team keeps the copyright. Uploading a drawing set to a service that stores it, reviews it, and may learn from it is a reproduction for a purpose nobody licensed. AIA's own help center cites a case where a contractor and its subs were held to infringe after the license ended. That is not a hypothetical risk category.
Then the rest of the pile. Subcontractor pricing is a trade secret only if you treat it like one, and the federal statute asks whether you took "reasonable measures" to keep it secret. Payroll, I-9s, drug tests, and injury reports are employee data with their own state rules. Site photos have faces in them. Anything from a federal or defense job carries markings that turn a paste into a reportable incident with a 72-hour clock. And most data-center, pharma, and healthcare owners now add a project-wide NDA on top of A201 with "no third-party service" language and audit rights.
None of that data is the AI vendor's problem. All of it is yours.
The gap is the tier, not the tool
The same product name can mean two very different contracts. The consumer tiers of the major assistants either train on what you type by default or require you to choose, and several keep it far longer than people assume.
| Consumer tier | Training on your inputs | Retention worth knowing about |
|---|---|---|
| ChatGPT Free and Plus | On by default, can be turned off | Deleted chats held up to 30 days |
| Gemini, personal account | On by default | Chats pulled for human review are kept up to three years, even if you delete them |
| Copilot, personal account | On unless you opt out or use a work account | Opt-out lives in account privacy settings |
| Claude Free, Pro, Max | You must choose | Five years if you allow training, 30 days if not |
| DeepSeek app | On, with an opt-out | Stored in the People's Republic of China; barred on DoD work |
The business tiers of the same vendors say the opposite in writing. OpenAI's API and Enterprise plans do not train on customer content. Anthropic's commercial terms say "Anthropic may not train models on Customer Content." Microsoft 365 Copilot on a work account keeps prompts inside your tenant. Google Workspace with a qualifying edition does the same. The difference between the two columns is a license fee and a signed data-processing addendum, and it is the whole story.
What "deleted" means also changed in 2025. In the New York Times case against OpenAI, a federal magistrate ordered OpenAI to preserve output logs that users had deleted. The order was narrowed months later, but everything captured in the window stayed on legal hold. OpenAI's own notice said Enterprise customers and API customers with zero data retention were not covered. Put plainly: data sitting in a vendor's system can be frozen by someone else's lawsuit, and the only way not to be in that pile is a contract tier that never stored it or a machine you own.
The leaks you did not sign up for
The paste is the obvious path. The quiet path is the AI that got turned on inside tools you already pay for.
A meeting note-taker joins the owner's OAC call because one attendee installed it. Otter's policy says it trains on de-identified audio and transcripts, and in August 2026 a federal court let wiretap and biometric claims against it proceed because the service "independently collects, retains, and uses communications for its own commercial purposes." HubSpot's AI training toggles default to on. Slack's global model training needs an email to opt out. A PDF tool's new connector hands the text of your drawings to whatever chatbot account the user has connected, which may be the free one.
Even the platforms built for this industry deserve a read. Procore publishes which models it routes to and says no third-party training happens on your data. Autodesk, Bentley, Trimble, and Oracle say similar things in their own words. Read the page rather than the sales deck, and ask the one question that matters: does my project data leave your cloud, and to whom.
There is also the phone. A superintendent photographs a sheet and asks a personal assistant app what a symbol means. That path bypasses every control on the laptop, and the fix is a work profile with copy-out blocked, not a memo.
The numbers say it is already happening
Construction has not had its Samsung moment in public. In 2023, engineers at Samsung pasted source code and a meeting transcript into a consumer chatbot three times in twenty days. The company banned the tools and built its own. No builder has made that headline yet. The surveys say the behavior is already here.
AGC's 2026 outlook, with 951 firms answering, found 61 percent use AI or plan to invest in it, and 45 percent already use it for office and administrative work, which is where the confidential documents live. Across industries, UpGuard found more than 80 percent of workers use unauthorized AI tools and executives use them most. IBM's 2026 breach study, released in July, put the global average cost of a breach at a record $4.99 million, found shadow AI involved in 43 percent of breached organizations, up from 20 percent a year earlier, and found 68 percent of breached organizations had no governance to detect unsanctioned AI use.
Insurers noticed first. Cyber applications in 2026 ask for an inventory of AI tools, an acceptable-use policy, and evidence of controls. Some carriers have added AI exclusions to professional liability forms. A renewal form is where many contractors will first be asked to describe their AI policy, and "we do not have one" is now an underwriting answer.
The other option: own the machine
Two years ago, running a capable language model in your own building meant a server rack and a specialist. In 2026 it means a desktop.
Think of private AI as a ladder with four rungs. At the top, open-weight models running on hardware you own, with the network egress blocked: nothing leaves the building, and the only outbound traffic is a model download. One rung down, single-tenant private cloud, where your prompts go to a region of a hyperscaler under a contract that forbids training and, on Azure, lets you turn off the 30-day abuse review. Below that, the enterprise tiers of the assistants, which promise no training but store what you type. At the bottom, the consumer tiers. Each rung down trades custody for capability and convenience.
The top rung got cheap. A 128 GB unified-memory desktop built on AMD's Ryzen AI Max+ 395, around $2,200 for a mini PC in mid-2026, runs OpenAI's open-weight gpt-oss-120b at roughly 30 tokens per second, fast enough for two or three people at once. NVIDIA's DGX Spark, about $4,000 to $4,700, runs the same model faster with the CUDA tools most integrators know. Apple's Mac Studio with the M5 Ultra, announced in August 2026 from $5,499 and configurable to 512 GB later this year, holds models in the 235-billion-parameter class. A workstation with one RTX PRO 6000 card, around $13,000 for the card alone, serves a 70-billion-parameter model to a whole office. All of it is exposed to this year's memory shortage, so quote hardware the week you buy it.
The models are open too. gpt-oss, the Qwen3 family including the vision models that read drawings and site photos, Mistral Small, DeepSeek, and GLM ship under Apache 2.0 or MIT licenses. Llama and Gemma work but carry custom terms worth reading. What these models do well is exactly the trailer's paperwork: answering questions across a project folder, summarizing a 400-page project manual, drafting an RFI, cleaning up meeting notes, and reading a photo of a sheet.
A sensible setup for a 10-to-50-person firm is one of those desktops on a UPS in the server closet, a second identical unit as a cold spare, and a short software stack: a document parser, an embedding model, a vector store that carries the same permissions as your SharePoint folders, a model server, a gateway that logs every prompt, and a chat window behind your single sign-on. The permission piece is the one to insist on. A PM should not be able to ask the assistant about another job's pricing, and that rule has to live at retrieval time, not in a policy document.
Costs, as estimates: a 50-person firm can stand this up for $10,000 to $60,000 all-in and run it for well under $1,500 a month including a fractional administrator. A 500-person firm should budget $100,000 to $250,000 to build and $4,000 to $12,000 a month to run, and most of that is people, not hardware. Compare it with the seat licenses you would otherwise buy, and with the cost of the one breach the IBM study prices.
Where the local box falls short
Open-weight models trail the frontier by roughly five to fifteen points on the hardest reasoning benchmarks. On a subtle contract interpretation or a long synthesis across many documents, a frontier model still wins. A local box also needs someone to patch it, back up the index, and swap models without breaking what worked last month. Private AI projects fail on operations, not on hardware.
The pattern that works for most contractors is a hybrid, routed by data classification rather than by task. Bid pricing, sub quotes, payroll, owner financials, incident narratives, and anything under an NDA stay on the local model over the local index. Public code questions, marketing copy, general research, and cleanup of non-confidential notes go to an enterprise API with zero data retention. One gateway fronts both, so people see one chat window and the auditor sees one log with the routing decision beside each prompt.
Whichever rung you choose, the verification steps are the same. Read the vendor's data-processing page. Confirm the retention setting in the admin console, not the brochure. Watch the egress logs for a week. Rehearse the incident.
The controls that cost nothing
Most of the protection is paperwork and settings, and it fits on a page.
The acceptable-use policy, in the language a crew reads: use only the approved tools, signed in with the company account. Never put "never paste" data into any AI tool, approved or not. AI output is a draft until a person with the right authority checks it. No note-taker joins a meeting unless the host announced it and outside participants were told. If something sensitive went into the wrong tool, tell IT the same day; nobody gets in trouble for reporting, people do for hiding it. Subs and vendors who touch our project data follow the same rules.
The classification, in three tiers. Never paste: controlled federal data, health information, owner-designated confidential information, unopened or unleveled bids, employee identity and medical records, security drawings, anything under litigation hold. Internal tools only: drawings and the project manual, schedules, change-order narratives, cost reports, photos, incident narratives with names removed. Any tool: public code questions, generic email polish with no project identifiers, training content. The field test a superintendent can apply is simple. Would I be comfortable if this showed up on the owner's desk with a note saying I sent it to a company we have no contract with?
The settings. Label confidential libraries in SharePoint or your project platform. Turn on the endpoint rule that blocks paste and upload of labeled content to the generative AI site group, and warns on the rest. Point the company's chatbot domain at your enterprise workspace only. Put every approved tool behind single sign-on. Manage the phones. Turn off "discoverable" sharing anywhere it exists; in 2025, roughly 4,500 shared chatbot conversations ended up indexed by Google because a checkbox existed.
The incident playbook, six steps. Contain the same day: delete the conversation, turn off training on the account, request deletion in writing and keep the screenshot. Classify on day one, because controlled federal data starts a 72-hour clock and an owner NDA usually has its own notice window. Notify per the contract and the policy, with counsel before any statement. Preserve the logs and the device. Fix the control that let it happen. Write one paragraph in the incident log, because the insurer will ask for it at renewal.
And the contract. Flow the same AI rules down to your subs, because their paste is your breach under the owner's clause. Watch for new owner language that bans third-party AI services outright, which would sweep in the platforms you already run, and negotiate it to "no service that trains on or retains project data outside a written agreement."
Custody
The instinct in this industry is to bolt the new thing onto the old process and see what happens. That works for a layout robot. It does not work for a tool whose entire value is that you hand it your documents.
The firms that will be fine are not the ones that avoid AI. They are the ones that decided, early and on paper, where their copies live. For some that is an enterprise contract with retention turned off. For a growing number it is a black box in the server closet with a hard hat hung on it, running a model they downloaded once and never phoned home again. Either way, the question to ask every tool, every vendor, and every new hire is the same one a super asks about a gang box at the end of the day. Who has the key.
Disclosure: this publication is run by SubPro, which processes customer project documents for submittal work under a written policy of deletion after delivery and no training on customer files. Nothing in this article recommends a specific vendor, including us.