# This Is The Tier One Application Nobody Could Find: The Connector List Alone Is Twenty Bits, So Compute Locally And Submit Banded

**version** v0.33.70
**date** 12 September 2026
**from** Human (project lead)
**to** Whoever builds the collector, whoever writes its schema, and whoever decides what leaves the browser

**type** Dev brief (specification for the shape collector, its schema, its submission path and its two modes)

*Second of 12 September. Both the corpus and the outside were searched. The corpus supplied a parked item this closes, because the ten pound product has depended since 10 September on an application built in August that nobody has located, and this memo respecifies it. The outside search was aimed at three things and changed the design on all three. On completion it found that the folklore about progress bars is backwards, with one indicator pattern raising drop off odds by more than half. On schemas it found, unusually for this corpus, that nothing already exists, because every standard describes the artefact and none describes the deployed configuration. And on collection it found that the memo's central assumption is wrong in a way that is measurable rather than arguable, because the connector question alone carries about twenty bits of entropy. Limitations: three of the completion studies are paywalled and are cited through an open access meta analysis; one vendor completion figure is marketing with no methodology and is treated as such; and the entropy arithmetic is an order of magnitude estimate.*

---

## What This Is

The specification for the interactive collector that stands in front of everything else, and the correction of one sentence in the memo that would otherwise be expensive: **the memo proposes a user interface that lets somebody interactively say what they have, with a feedback loop, questions on one side and a picture accumulating on the other, asking which models they use, on which surfaces, and most importantly which connectors are enabled on each, then representing that back to them visually, on the reasoning that this is the beginning of the workflow and that from it a first version of a policy or at least an ingredient list can already be produced; it proposes a standardised schema with an ontology and taxonomy behind it mapping the permission settings, a vault that collects the responses and, more interestingly, a vault that submits into another vault so that the multi vault capability is demonstrated; it observes that responses could be collected anonymously and then says that they can be captured anyway because it is all anonymous; it proposes designing the whole thing as a game with a game like feel; and it identifies a second product, which is the same instrument sent by a company to its own staff to find out how many agents are in use, potentially branded and sold; the first finding is that this closes a parked item, because it is the tier one deliverable that has been specified since 10 September and blocked on a missing application, and building it as the collector makes the ten pound product real; the second is that the sentence about capturing it anyway is wrong and the error is measurable, because a twenty way connector multi select carries about twenty bits, the nominal configuration space runs to the order of a trillion cells, and at five hundred responses effectively every respondent is unique, which puts the data squarely inside the definition of personal rather than outside it; the third is that the fix is cheap and keeps the product intact, being that the full configuration is computed in the browser and never leaves, while what is submitted is a banded version with connector counts and categories rather than a list, which is the largest single reduction in entropy available; the fourth is that the write only lane between vaults that the memo is excited about is the right mechanism for exactly this reason, because a collector that cannot read back cannot correlate; the fifth is that designing it as a game is the wrong frame, because this is a data collection instrument rather than a teaching one, the evidence that matters is survey completion rather than serious games, and in that literature a constant progress bar has no effect while one that visibly stalls raises drop off odds by fifty six per cent; and the sixth is that the employer mode is a different product under different law, because consent is a weak basis in employment, an assessment is likely required before it runs, and the employer holds exactly the identifying context we do not.** New contributions: **the identification of this as the missing tier one application; the entropy arithmetic and the banding that fixes it; the local computation and write only submission design; the completion evidence and what it says about the game framing; the finding that no schema exists for a deployed configuration, which is a gap rather than a thing to reuse; and the separation of the employer mode with its own rules.**

## This Is The Application Nobody Could Find

**On 10 September the tier one product was specified as their own answers, their measured grant and the delta, as a file they keep, at ten pounds.** It has carried a dependency ever since: an application built in August that nobody has located, without which tier one is a questions page.

**This memo is that application, described from scratch and better.** The collector asks the shape, computes the grant from it, and returns the delta. **That is tier one, and it needs no August artefact.**

**And it is the first four stages of the flow specified yesterday**, which were the only ones that could ship this week: encounter, self select, draft, correct. **The collector is stages two and three, and the correction is stage four happening inside the same screen rather than in a conversation.**

**So the parked item closes and the ten pound product stops being conditional.** That is the most useful thing in this brief and it should be recorded as such.

**On the naming.** The memo says the output might be an ingredient list and needs a name. **It has one: the shape.** The corpus has used deployment shape since 11 September for exactly this, the collector produces a shape record, and the grant is derived from it. **No new noun, which is the ruling of 4 September.**

## The Sentence That Has To Go

> Actually, we don't even need that because I can just capture it anyway. So I can, because to be honest, it's all anonymous.

**It is not anonymous, and this is measurable rather than arguable.**

**The arithmetic.** A multi select over twenty connectors is two to the twentieth, which is about a million combinations, or **roughly twenty bits**. Add assistants, surfaces, role and company size and the nominal space runs to the order of **five trillion cells**. Real distributions are skewed, so far fewer cells are occupied, but the comparison that matters is with the number of respondents:

| Responses | Cells needed for a group of five | Realistic outcome |
|---|---|---|
| 100 | at most 20 | **Effectively every respondent unique** |
| 500 | at most 100 | **Effectively every respondent unique** |
| 5,000 | at most 1,000 | Common shapes reach a group of five. **Every interesting shape does not** |
| 50,000 | at most 10,000 | The head is safe and the tail still singles out |

**Two anchors confirm it.** Three low cardinality fields, being postcode, gender and date of birth, uniquely identify about **eighty seven per cent** of a national population. A browser fingerprint study of four hundred and seventy thousand samples found **eighty three point six per cent instantaneously unique** at about eighteen bits. **Our connector question alone carries more entropy than that study's whole fingerprint.**

**The regulator's own test is singling out**, meaning the ability to isolate the records relating to one person, and the assessment is against a motivated intruder assumed reasonably competent with ordinary resources. **A configuration fingerprint is a fingerprint in the literal sense**, and the regulator cites a standard of groups of five as strong protection, which nothing above reaches.

**So submissions are personal data and must be processed as such.** That is not a reason not to build it. It is a reason to build it the way described next, which costs almost nothing and makes the product better.

## Compute Locally In Full, Submit Banded

**The resolution keeps everything the product needs and gives up nothing the user wants.**

**In the browser, with full detail.** Every connector, every surface, every assistant. The grant is computed from the full shape, the delta is derived, the label and the leaflet render. **The user sees everything. It is their configuration and it never leaves their machine.**

**What is submitted, and only if they press submit**, is a banded record:

| Field | Full, local | Submitted |
|---|---|---|
| Connectors | The named list of twenty | **A count band, being none, one to two, three to five, six or more, plus categories: mail, files, calendar, code, chat** |
| Assistants | The named list | A count band and the categories |
| Surfaces | Each one | Coarse, being desktop, web, editor, command line |
| Role | Free choice | **Three or four bands** |
| Company size | Exact | **Three bands** |
| Anything typed freely | Shown back to them | **Never submitted.** Free text cannot be anonymised |
| Address, user agent, precise time | Not needed | **Never recorded**, since each one independently defeats the banding |

**The connector banding is the single largest entropy reduction available**, and it costs the research almost nothing, because the interesting finding is *how many* and *of what kind*, not which brand.

**Suppress on output.** Never render a cell backed by fewer than five submissions, and check that a suppressed cell cannot be recovered by subtracting the ones around it.

**And write down the assessment**, naming who the motivated intruder would be. The obligation is not to reduce the risk to zero; it is to make a reasoned assessment and record it.

## The Write Only Lane Is The Right Mechanism, And Now There Is A Reason

**The memo is excited about a vault that submits into another vault, as a demonstration of the platform. It is more than a demonstration, it is the correct privacy architecture for this, and that is worth saying because it turns a capability into an argument.**

**A write only lane means the collector cannot read back.** It cannot correlate a submission with an earlier one, cannot join across respondents at the point of collection, and cannot reconstruct a session. **A collector that is structurally unable to correlate is a stronger claim than a collector that promises not to.**

**That is the estate's own enforcer test applied to our own product.** A promise not to correlate is bounded only when enforced by something the collector does not have. **Here the thing it does not have is read access.**

**One rule that follows and must be in the build.** The submission carries **no stable identifier of any kind**. Not a session token, not an install identifier, not a hash of anything. The pattern to copy is the distribution install counter noted on 10 September, which refuses identifiers explicitly and buckets by week.

## Two Regimes, And Only One Of Them Needs Consent

**The legal analysis splits cleanly and the design should follow the split.**

| Operation | Regime | Position |
|---|---|---|
| Storing the in progress answers in the browser so the tool works | Storage rules, in scope | **Strictly necessary.** Assessed from the user's point of view, and somebody who starts a self assessment plainly wants their answers to persist. **No banner** |
| Reading them back to render the result | Storage rules, in scope | Same exception, same reasoning |
| Analytics about the tool itself: drop off point, device class, which question loses people | Storage rules, in scope | **The statistical purposes exception fits**, with clear information and a simple free means to object, which the regulator suggests as a toggle defaulted on. **Browser settings are not sufficient** |
| Submitting the banded answers | Not the storage rules. Data protection only | **An explicit, separate, unticked act.** The answers are the content the user produced, not statistics about how the service is used, and the regulator is explicit that the exception is not a broad one covering all analytics |

**That fourth row is the precise point the memo gets wrong.** The exception covers **how** the service is used. It does not cover **what the user told it**. Those are different things and only the first is exempt.

## Do Not Design It As A Game. Design It To Finish

**The memo's instinct is that the game framing carries over. It does not, and the evidence that applies here is a different literature.**

**This is a data collection instrument, not a teaching artefact.** The serious games evidence of 10 September is about learning outcomes. **What matters here is completion**, and that literature is well developed and mostly contradicts the folklore.

**The progress bar finding is the sharpest and it is backwards from what everybody believes.** A meta analysis covering nineteen studies and thirty two experiments found:

| Indicator | Effect on dropping out |
|---|---|
| **Constant speed, an honest linear bar** | **No significant effect** |
| Fast at first then slowing | Drop off odds multiplied by about **0.80** |
| **Slow at first then speeding up** | **Drop off odds up by fifty six per cent** |

Its own conclusion is that the findings question the common belief that progress indicators reduce drop off. **And the harmful pattern is exactly what this tool would produce by accident**: a short introductory section followed by a long connector by connector section makes the bar stall in the middle.

**A second finding kills the other assumption.** One question per screen does not reliably buy completion. A controlled comparison of a conversational form against an ordinary one found completion of **eighty nine point eight against ninety one point four per cent**, no significant difference, and the conversational version took **fifty three per cent longer**. A 2026 field experiment found it rated more original and entertaining, harder to navigate, and with no practical data quality advantage.

**And the one widely quoted commercial figure does not survive inspection.** A conversational form vendor claims forty seven point three per cent completion against a stated industry average of twenty one and a half. **No methodology, no denominator definition and no independent verification**, and its rich media claim of plus a hundred and twenty per cent is an uncontrolled correlation between completion and how much effort somebody put into the form.

**What actually works, ranked by evidence quality.**

| Rank | Lever | Evidence |
|---|---|---|
| **1** | **Fewer questions** | The only lever with both randomised and large observational support, and the observational data shows a cliff above roughly fifteen |
| **2** | **An honest, short stated duration** | Randomised and replicated. **The announcement is itself a treatment**, and a progress indicator only helped when the task was promised short and actually was |
| **3** | **Works on a phone, one item at a time, no grids** | Strong on quality, good on completion. Most responses will be mobile |
| **4** | **A high value question early** | One clean experiment, one large observational study. Modest but real |
| **5** | Showing partial results | **Satisfaction significantly higher, completion roughly unchanged.** Treat as an untested hypothesis |

**So the memo's accumulating picture on the right is worth building, and for the honest reason: it makes people feel the thing is worth finishing and it makes the output legible. It is not evidenced as a completion mechanism and should be measured rather than assumed.**

**The one game mechanic that earns its place is the estate's own.** Before revealing what the selected connectors grant, ask the user to guess the number. **That is the published mechanic, it costs one screen, and it is what makes the result land rather than scroll past.** It is also the thing that turns a form into the experience the memo is reaching for, without the fifty three per cent time penalty of a conversational interface.

**Concretely: target under fifteen questions, say honestly how long it takes, one item per screen on mobile, a front loaded progress indicator or none at all, and one guess screen before the reveal.**

## Nothing Like This Schema Exists, Which Is Unusual Here

**Twice this month the answer to do we need to build this has been that it already exists. Not this time.**

**Every standard in the area describes the artefact rather than the deployment.** The component bill of materials for machine learning, standardised as an international specification in December 2025, describes a model: its parameters, its task, its architecture family, its datasets, its inputs and outputs, its quantitative analysis and its considerations. The other bill of materials family has an equivalent profile. **None of them describes which assistant a person runs, on which surface, with which connectors granted which scopes.**

**So the schema the memo asks for is a genuine gap and worth building.** With three disciplines the corpus already has.

**Reuse the capability grammar.** The shape's whole purpose is to compute a grant, and the grant is expressed in the published primitives of verb, object and reach. **The schema's job is to map a configuration to a set of those, and nothing more.**

**Derive rather than assert.** The mapping from a connector to the capabilities it grants is the researched material from yesterday, and every row carries a source, a date, and whether it was measured or derived.

**Publish it as data in the open repository**, so that a correction is a proposal with evidence attached, which is the four layer architecture of 10 September.

**And do not call it an ontology until it needs to be one.** The finding of 10 September stands: the value is in a closed controlled vocabulary, not in the formalism around it.

## The Employer Mode Is A Different Product

**The memo's second idea is a real product and it is not the same product, because the law changes when the employer sends the link.**

**Consent stops working.** The regulator's guidance on monitoring workers says consent is not usually appropriate in the employment context because of the imbalance of power, and that it must be freely given and withdrawable without detriment. **A survey circulated by an employer, whose answers could reveal a policy breach, is close to the worst case for freely given consent.** The realistic basis is legitimate interests with a documented assessment.

**An impact assessment is likely required.** The regulator says one must be carried out before processing likely to cause high risk to workers' interests, and a tool that inventories which systems an individual employee uses, on which devices, with which data connectors, is systematic monitoring of workers.

**And this is the part that matters most for the design.** Aggregated results about staff are personal data in the employer's hands even when they are not in ours. **The regulator recognises relative anonymity explicitly: information can be personal data in the hands of one organisation and anonymous in the hands of another that lacks the context.** The employer has the organisation chart, the device fleet and the sign on logs. **A row reading one product person on a desktop assistant with mail, files and code connectors is anonymous to us and a name to them.**

**So the employer mode never returns an individual row.** Only cells above the suppression threshold, only bands, and the product says so on the page the employee sees before they answer. **That is also the only version an employee will answer honestly, which makes the privacy constraint and the data quality constraint the same constraint.**

## Nobody Else Asks. That Is The Opening And The Problem

**A survey of the discovery market found no vendor using self report.** Every one of them discovers through technical telemetry: browser extensions, endpoint agents, network inspection, application programming interface integrations, sign on logs. One free assessment tool disparages self report explicitly, recommending real usage telemetry instead.

**Both halves of that matter.**

**The opening.** Self report is the only method that reaches a personal device, a personal account, and a locally run server without deploying anything. One of the larger vendors implicitly concedes the gap by selling desktop agents to close it. **And it needs no procurement, no installation and no administrator.**

**The problem.** No incumbent treats survey data as a legitimate discovery source, so there is no established vocabulary of trust for it. **The honest framing is that this finds what people will tell you, which is different from what is on the network, and neither is complete.** That belongs on the page rather than in a footnote.

**The closest comparator is a free seven question organisational assessment taking about two minutes with no email gate.** That is the completion standard to beat. **And nothing found asks an individual which assistants and connectors they personally use and returns them a result**, which is the specific thing being built.

## What This Does Not Try To Be

- **A schema.** The gap is established, the disciplines are named, and no field is defined.
- **A design.** Question count, ordering and the accumulating picture are constrained by evidence. Nothing is drawn.
- **Legal advice.** The split between the two regimes and the employment analysis are research. An impact assessment for the employer mode is a piece of work, not a paragraph.
- **A privacy claim.** The word anonymous should not appear in the product until the banding, the suppression and the assessment all exist. Until then the honest word is banded.
- **A completion promise.** The evidence constrains the design and predicts nothing about our numbers.

## Honest Tensions

| Tension | Note |
|---|---|
| Banding the submission | It is what makes the data lawful to hold, and it throws away the brand level detail that would be most interesting to publish |
| Computing locally | It is the strongest privacy position available, and it means we learn nothing unless somebody presses submit |
| Not designing it as a game | The evidence is about completion rather than fun, and the memo's instinct about feel is what will make people start |
| The guess screen | It is the one mechanic with a published argument behind it, and it adds a screen to an instrument whose main lever is fewer screens |
| Employer mode | It is a real second product, and it carries an impact assessment, a weak lawful basis and an intruder who is the buyer |
| Self report | It reaches what telemetry cannot, and no incumbent treats it as legitimate |
| Building the schema | It is a genuine gap, and it is the third schema this estate has taken on in a fortnight |

## Open Questions

1. **How many questions, exactly?** The cliff is around fifteen and the connector question alone could be twenty items if built carelessly.
2. **Is the connector question one multi select or a category then a count?** The second is the banding made native, and it may lose people who want to see their tool named.
3. **What does the guess screen ask?** One number, or one number per category.
4. **Who writes the motivated intruder assessment, and where does it live?** It is a short document and it is a precondition for using the word anonymous.
5. **Does the employer mode need a different instrument or the same one with a different output?** Same instrument is cheaper and the consent problem does not care.
6. **What is the suppression threshold?** Five is the cited standard and a small early sample will suppress nearly everything.
7. **Does the shape record schema become the published vocabulary, or stay internal until the fifth policy?** Publishing early invites correction and locks a shape too soon.

## Relationship To Previous Briefs

**From the payment brief of 10 September**, it takes the tier one specification and closes the dependency that has blocked it since.

**From the end to end brief of 11 September**, it takes stages two, three and four of the flow, and the researched connector grants that the schema maps.

**From the toolkit brief of 10 September**, it takes the ruling that the transmitting code is absent rather than disabled, the storage regulation analysis, and the install counter that refuses identifiers.

**From the teaching brief of 10 September**, it takes the games evidence and finds that it does not apply here, because this instrument teaches nothing and must only be finished.

**From the games site**, it takes the one mechanic that does apply, which is stating a belief before being shown the answer.

**From the findability brief of 10 September**, it takes the rule that the value is a closed vocabulary rather than a formalism, and applies it to a schema that genuinely has to be written.

**From the vault architecture brief of 10 September**, it takes the four layers, and places the shape schema in the community editable data layer.

**From the enforcer test of 20 August**, it takes the form of the argument for the write only lane: a promise not to correlate is bounded only when enforced by something the collector does not have.

## Key Claims

| # | Claim |
|---|-------|
| 1 | This is the tier one application that has been missing since 10 September, so the ten pound product stops being conditional |
| 2 | The output already has a name, which is the shape, and no new noun is needed |
| 3 | A twenty way connector question carries about twenty bits, which is more entropy than a well known browser fingerprint study measured in a whole fingerprint |
| 4 | At five hundred responses effectively every respondent is unique, so the submissions are personal data |
| 5 | The fix is to compute the full shape locally and submit only bands and categories, which is the largest entropy reduction available and costs the research almost nothing |
| 6 | A write only lane cannot correlate, which is a structural claim rather than a promise, and it needs no stable identifier of any kind |
| 7 | Storing answers to make the tool work is strictly necessary, and submitting them is a separate explicit act under a different regime |
| 8 | The statistical exception covers how a service is used and not what a user tells it, which is the precise point the memo gets wrong |
| 9 | A constant progress bar has no measured effect and one that stalls raises drop off odds by fifty six per cent |
| 10 | A conversational form showed no completion advantage and took fifty three per cent longer, so fewer questions and an honest duration are the levers |
| 11 | No existing standard describes a deployed configuration, so this schema is a genuine gap rather than something to reuse |
| 12 | In employer mode the employer holds the identifying context we lack, so the output is bands and suppressed cells and never an individual row |

---

## Sources

All read 12 September 2026.

**Inside the estate.** The capability map and its twenty three primitives at https://what-can-it-do.games.sgit.ai/map/index.html. The games site for the belief before answer mechanic at https://games.sgit.ai/llms.txt. The payment, toolkit, teaching, findability and vault architecture briefs of 10 September, and the end to end brief of 11 September.

**Completion.** The progress indicator meta analysis covering nineteen studies and thirty two experiments, Villar, Callegaro and Yang 2013, Social Science Computer Review 31(6), which reports no significant effect for a constant indicator and a fifty six per cent increase in drop off odds for a slow to fast one. The expectation matching result, Yan, Conrad, Tourangeau and Couper 2011, International Journal of Public Opinion Research 23(2). The conversational comparison reporting ninety one point four against eighty nine point eight per cent completion and a fifty three per cent time penalty, Kim, Lee and Gweon 2019, CHI. The 2026 field experiment finding engagement up and no practical quality gain, Cavusoglu Deveci, Fuchs and Metzler, Bulletin de Methodologie Sociologique 169-170(1), published 6 March 2026. The item by item finding, Revilla, Toninelli and Ochoa 2015. The personalised feedback trial finding satisfaction up and response behaviour largely unchanged, Kuhne and Kroh 2018, Social Science Computer Review 36(6). The vendor completion claims at https://www.typeform.com and the January 2024 release behind them, treated as marketing with no methodology.

**Schemas.** The component bill of materials for machine learning, version 1.7, standardised as an international specification in December 2025, at https://cyclonedx.org/. The alternative bill of materials family's profile at https://spdx.github.io/spdx-spec/. Neither describes a deployed configuration.

**Identifiability.** The regulator's guidance on effective anonymisation, its singling out test, its motivated intruder test and its citation of a group of five as strong protection, at https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/how-do-we-ensure-anonymisation-is-effective/. The finding that three low cardinality fields uniquely identify about eighty seven per cent of a population, Sweeney. The browser uniqueness study of four hundred and seventy thousand samples finding eighty three point six per cent instantaneously unique, Eckersley 2010.

**Storage and submission.** The guidance on storage and access technologies, published 29 April 2026, at https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/guidance-on-the-use-of-storage-and-access-technologies/, including the strictly necessary test assessed from the user's point of view, the statistical purposes conditions, the requirement for an objection mechanism that is not a browser setting, and the statement that the exception is not a broad one covering all analytics. The commencement of the new exceptions on 5 February 2026.

**Employment.** The regulator's guidance on monitoring workers, originally October 2023 and last updated 16 June 2026, at https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/employment/monitoring-workers/, including that consent is not usually appropriate because of the imbalance of power and that an assessment must be carried out before processing likely to cause high risk to workers.

**The discovery market.** A ten vendor comparison published 7 September 2026 stating that all ten rely on technical telemetry and none on self report. Vendor pages for browser, endpoint, network and integration based discovery, and the one transparent price found, at https://www.nudgesecurity.com. The free seven question organisational assessment at https://aona.ai/tools/shadow-ai-risk-assessment, which recommends telemetry over self reporting.

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
