← All research

How we assess this

GPT-5.5 partner pack (pressure-test only)

Executive summary

Verdict

Proceed, but change both the promise and the sequence. The codebase contains the beginning of a credible product, but it is not yet the product described by the most ambitious strategy documents. The defensible proposition is not “speak into a phone and recover a claim.” It is:

Turn fragmented field and project evidence into source-grounded facts; reconcile contradictions across media; surface the few evidence gaps that are about to become costly; route them to an owner; and only then project the record into operational, commercial, contractual, or handover workflows.

That is a useful, differentiated control layer. A voice-to-diary application, an AI dashboard, a generic contract chatbot, or a claims pack assembled from fictional facts is not.

The market research correctly identifies a seam between upstream evidence capture and downstream commercial judgement, but it overstates the maturity of the moat. Field capture is crowded and becoming inexpensive; claims and recovery products do exist; cross-firm benchmarking remains empty largely because permission and trust are hard, not because nobody has noticed the opportunity. The synthesis itself records these corrections at research/dossiers/_LANDSCAPE-SYNTHESIS.md:L23-L44. The current advantage is therefore a workflow and trust thesis, not a data moat already earned.

The seven decisions

1. Keep the thesis, but move the centre of gravity

Keep the evidence-to-commercial-event concept, but do not make CommercialEvent the first thing the model claims to know. The sequence should be:

source material
  → observations with exact provenance
  → facts with explicit state
  → cross-source reconciliation
  → evidence gaps, owners and decay
  → event hypotheses
  → contract/jurisdiction projection
  → human-reviewed action, record or recovery pack
  → outcome label

The strongest implementation already points this way. apps/api/evidence_analysis.go:L8-L162 defines bounded analysis outputs, abstention, field-level states, geometry and trace metadata. apps/api/evidence_reconcile.go:L61-L213, :L873-L970, :L1017-L1148 and :L1352-L1392 distinguish corroboration, conflict, unreadability, wrong-document, uncomparable evidence and non-analysis, then preserve source regions and route specific remediation. That is considerably more credible than the recovery surface, which still relies on canned profiles and scenario mappings in apps/api/profile.go:L40-L274 and apps/api/pipeline.go:L287-L330.

2. Rebuild the customer-facing demo around one trust decision

For Mountain Earth, the single impressive moment is not the A$29,000 recovery story. It is this:

  1. A site manager describes a live M&E delivery or access issue in 60–90 seconds and attaches a delivery docket and photo.
  2. The system extracts a reported fact, reads the docket, and either corroborates it, contradicts it, or abstains.
  3. The reviewer clicks the fact and sees the exact transcript span, docket text or image region.
  4. One evidence gap with a real expiry window is assigned to the right person; when the missing source is supplied, the record resolves.
  5. A Procore-ready daily log/change-event/RFI support record is generated, with no invented entitlement or quantum.

Frame it as: “Procore records what people enter. We are showing the control before a record becomes trustworthy.” Mountain Earth publicly says it uses Procore for safety tracking and cross-site visibility, while its own discovery call described Procore as a roughly £25,000-per-year system spanning design, commercial, project, quality and safety workflows (research/_source-conversation/turn-01-user-will-help-me-think-through-this.md:L33-L36). Competing with that system of record would be strategically incoherent.

Cut the NSW/AUD/AS4000/Security-of-Payment framing, the injury-as-commercial-proof beat, the nine-step product tour, object counting as progress proof, and any statement that the fictional project “recovered” money. The current story is explicitly fictional and Australian (demo/scenarios/riverside/storyline.md:L1-L30); its late-delivery-to-injury-to-A$38,000-variation-to-A$29,000-shortfall chain (:L1-L16) will read as theatrical and geographically inattentive in a UK principal-contractor room.

3. Ask for a closed matter, not a data lake

The highest-value ask is one closed project with a known commercial outcome, selected because somebody in the room can explain what happened and why. Request, in order:

  1. outcome truth and a guided matter walkthrough;
  2. executed contracts and amendments;
  3. change, payment and final-account truth;
  4. correspondence, RFIs, instructions and meeting records;
  5. diaries, dayworks, labour, delivery and cost proof;
  6. native programme files and updates;
  7. drawings, specifications, revision and submittal history;
  8. photos, inspections and pre-cover evidence;
  9. Procore workflow/configuration/export samples.

Do not begin by asking for every historical file. Begin with a manifest, a narrow date/event window and a read-only data room under NDA/DPA. Exclude privileged material, pseudonymise parties, prohibit raw-customer-data training by default, define retention/deletion, and reserve cross-firm derived use for a separate explicit opt-in. The existing partner questions correctly ask for a painful project and its contract, payments, diaries, correspondence, drawings, programme, resources and outcome (research/06-build-spec/output/70-open-questions-for-the-partner.md:L37-L60), but they need to be converted from an interview list into a labelled evaluation corpus with known outcomes.

4. Speak to their operating model, not to “construction” in the abstract

Mountain Earth is a North West principal contractor with integrated in-house M&E, a flat senior-led structure and a strong public emphasis on live environments, early coordination, accountability and transparent reporting. Its public projects include a £5m+ occupied HVAC replacement at 101 Old Hall Street, a £25m constrained residential scheme at Uptown, and a £7m multi-building University of Leeds programme. Its own prior call highlighted approximately 30% proper diary completion, the difficulty of reusing historical prices, the need for earlier director visibility as the business grows, and interest in live O&M, asset/material tracking, photo evidence and programme status (research/_source-conversation/turn-01-user-will-help-me-think-through-this.md:L49-L123).

This creates a specific positioning constraint. Mountain Earth publicly promotes fast, direct problem-solving and says its integrated model avoids defensive correspondence and paper-trail RFIs. A “claims bot” will sound like the bureaucracy it believes its model eliminates. Present the product as the memory and accountability layer that preserves a ten-minute decision without turning it into a three-day administrative process.

5. Lead with three capabilities; defer the seductive traps

Lead with:

  1. Multimodal reconciliation with provenance and abstention. This is the strongest code and the clearest trust demonstration.
  2. Evidence ingestion, expiry/gap routing and Procore-ready export. This fits the target’s existing stack and converts capture into an operational control.
  3. Project-specific UK contract/action assistant. Ingest the executed contract and amendments, calculate candidate clocks and required records, and keep every formal output draft-only until reviewed.

Treat STT, OCR and a very simple field interface as enabling infrastructure rather than the moat. Use CV first for constrained, auditable tasks: reading dockets; detecting whether a required object or label is visible; comparing repeated views under a defined capture protocol; linking regions to facts. Do not lead with arbitrary-photo measurement, autonomous progress percentages, generic drawing takeoff, safety detection, or fully autonomous notices/claims. apps/api/evidence_reconcile.go:L873-L923 already encodes the right discipline: counts are comparable only under matching model, prompt, threshold, ROI, viewpoint and capture protocol, and visible quantity is not installed quantity or completion.

6. Do not stay AU-first commercially; do not go globally generic

Adopt a UK design-partner wedge with a generic evidence/event kernel. Build JCT/NEC/HGCRA-oriented packs for the live customer, while retaining the Australian Security-of-Payment pack as a second reference implementation and regression test.

The current schemas mix universal semantics and local law. research/06-build-spec/output/schemas/event-taxonomy.au.yaml:L1-L48 combines generic event identity and evidence with AU contract hooks; :L529-L569 embeds AU payment-shortfall and adjudication routes. research/06-build-spec/output/schemas/common.defs.schema.json:L19-L99 hard-codes currencies, jurisdictions, contract forms and event types, while commercial_event.schema.json:L204-L235 puts AU recovery routes into the base schema. These should become registry IDs and versioned rule-pack references. Project timezone must be explicit rather than defaulting to Australia/Sydney, as occurs in apps/api/handlers_field_v1.go:L1295-L1315 and research/06-build-spec/output/schemas/project.schema.json:L57-L60.

7. Be product-first and service-assisted; never be free-form free consulting

Offer a fixed-scope paid design-partner engagement:

A sensible first fee is £15,000–£25,000 for the diagnostic/back-test, or £7,500–£15,000 at a documented design-partner discount that expires and is creditable against a pre-agreed annual subscription. Warm introductions and data access are valuable consideration but should not replace all cash. The product spec itself says “Never a free pilot” and proposes £9,000–£15,000 evidence audits and £15,000–£38,000 reconstruction engagements (apps/spec/field-record-design-spec.md:L274-L291).

Customer owns raw data and customer-specific records. You own the platform, schemas, prompts, agent skills, connectors, evaluation harness, generic workflow knowledge and benchmark methodology. Any de-identified aggregate use must be expressly granted. Avoid work-made-for-hire, broad exclusivity, roadmap veto, perpetual free licences, uncapped liability, recovery guarantees and unlimited custom integration.

What is strong now

What remains demo scaffolding

90-day decision plan

Days 0–30: make the demo honest and target-specific

Days 31–60: prove on history

Days 61–90: run one constrained live workflow

Go/no-go gates

Proceed to a live paid product pilot only if all are true:

  1. A customer gives access to one sufficiently complete closed matter under usable data/IP terms.
  2. At least one operational owner and one commercial/document-control owner attend the workflow design and weekly review.
  3. Provenance is correct on at least 95% of facts accepted by reviewers; unsupported claims are not silently promoted.
  4. High-value evidence gaps are recalled at an agreed threshold without creating an unmanageable false-positive queue.
  5. The workflow saves reviewer time or prevents a demonstrable evidence failure versus the current process.
  6. The customer agrees to a defined paid conversion and does not require broad exclusivity or ownership of core IP.

Stop or reposition if the only enthusiasm is for diary transcription, generic dashboards or bespoke consulting reports; if no complete historical matter can be supplied; if Procore plus existing discipline is already sufficient; or if every project requires a new ontology and extensive manual legal/QS interpretation.

Bottom line

This is recommended with high risk. There is a credible product inside the repository, and the target firm is unusually suitable as a design partner because it is technically led, operates live and constrained projects, already uses Procore, and has articulated the exact information failures being addressed. The opportunity will be lost, however, if the team mistakes a polished demo for a recovery engine, treats frontier-model access as a moat, or accepts a relationship in which Mountain Earth receives custom consulting and owns the resulting product. The next proof is not another feature. It is one closed project, one known outcome and an auditable comparison between the product’s reconstruction and the people who lived it.