# Collecting candidates

Collection produces candidates, not prospects. A candidate is an organisation that matches the shape you described, with almost nothing asserted about it yet — a name, a public presence, whatever identified it. No score, no fit reason, no address and no judgement. Keeping this stage judgement-free is what stops a cheap first impression from setting the tone for everything read afterwards.

- **Status:** Available
- **Audience:** both
- **In the app:** #/prospects
- **Last verified:** 2026-09-10
- **Canonical:** https://connectbyjbrh.com/docs/prospects/candidate-collection/

## What a candidate is, and is not

The vocabulary is load-bearing. A **candidate** is a name that matched. A **prospect** is an organisation that has been researched, has evidence attached and has been judged against your criteria. Almost every confusing prospect screen is somebody reading a candidate as though it were a prospect.

|  | Candidate | Prospect |
|---|---|---|
| Has a name and a public presence | Yes | Yes |
| Has evidence attached | No | Yes, or the claim is not there |
| Has a fit judgement | No | Yes, with a written reason |
| Has a sourced contact address | No | Sometimes — and only if one was found |
| Can be written to | No | Only after readiness and the compliance check |

This is also why the count of candidates is not a count of anything you can act on. A run that collects two hundred candidates and qualifies eleven has not failed; it has done the filtering you would otherwise do by hand, and the eleven are the number that matters.

## Where a candidate comes from

- A public source describing organisations of the kind your criteria name.
- A directory or listing where a business publishes its own trade and place.
- A reference from another researched organisation — a supplier, a partner, a named client.
- Your own list, imported rather than discovered — see [Uploading your own list](/docs/prospects/prospect-upload/).

Whatever the origin, the origin itself is kept. An imported row and a discovered one look different on the record on purpose, because the question *where did this come from* is the first one anybody asks about an unexpected prospect, and it is asked months later when nobody remembers the run.

## What is deliberately not decided yet

Three things are withheld at this stage, and each has a reason.

**No contact address** — Addresses come from research with a source recorded, never from the name of the organisation. Deriving one at collection time would be exactly the guessing that Connect does not do.
**No score** — Scoring before reading would rank candidates on how well they matched a filter, which is a fact about your wording rather than about the business.
**No fit reason** — A reason written before evidence exists is a sentence about a category. The reason is written last, from what was actually read — see [Why this prospect](/docs/prospects/fit-reason/).

One thing is decided immediately: whether this organisation is already yours. Identity resolution runs early precisely because it is cheap and because the cost of getting it wrong is high. A candidate that resolves to an existing person or company leaves the cold list before any research money is spent on it.

## Duplicates, near-duplicates and noise

Public sources describe the same business more than once — a trading name and a registered name, two branches, an old listing beside a current one. Collection does not silently merge them. Duplicates are **proposed**; merging is a human decision, and it preserves both sides' identities, stages, follow-ups, deals and cases rather than choosing a winner.

> **Careful** Resist merging early. Two entries that look like one business are sometimes two, and an unmerge is not a clean operation once outreach, replies and a pipeline entry have accumulated against the merged record.

Noise is different from duplication: a directory row that is not a business at all, or a business that closed. Those survive collection and die at research, when the sources that would have supported a claim about them do not exist. That is the intended division of labour — collection is broad and cheap, research is where the population gets honest.

## Questions

### Why do I see organisations that obviously do not fit?

Because collection matches the shape you described without judging it. The filtering happens at qualification, after reading. If the misfits share a property you could have excluded, add it as a disqualifier to the criteria and the next run will not collect them at all.

### Are candidates counted in the numbers above the list?

Organisations discovered counts them; contactable and qualified do not, because neither is known yet. Those three numbers being far apart is normal and is the clearest picture of where a run stands.

### Can I delete candidates I do not want?

You can, but excluding them in the criteria is usually better: deleting removes this run's rows, while a disqualifier stops them being collected again the next time the same search runs.

## Related

- [Describing who you want to reach](https://connectbyjbrh.com/docs/prospects/discovery-criteria/)
- [Researching a candidate](https://connectbyjbrh.com/docs/prospects/public-source-research/)
- [Resolving a prospect to an existing relationship](https://connectbyjbrh.com/docs/prospects/prospect-identity/)
- [Uploading your own list](https://connectbyjbrh.com/docs/prospects/prospect-upload/)
- [Prospecting in Connect](https://connectbyjbrh.com/docs/prospects/)
- [Relationships in Connect](https://connectbyjbrh.com/docs/relationships/)

## What this page is based on

- docs-source/sources/CHANNELS.md §4 — the stage order and identity resolution
- docs-source/sources/CHANNELS.md §5 — duplicates propose, merge_people is a human decision
- `docs-source/facts.py` — CANONICAL_TERMS Prospect, Person, Company
