Refinement

How a CSV is keyed

Every row of an upload is matched to the product and to the tracked companies on four keys: email | batch | name | project_ids. The header must have batch and at least one of email or project_ids; the other columns are kept as written.

The four keys

project_idsor project_id
The product projects the row names, directly. The cell may hold one id or several (12, 34); it is stored as the distinct ids, sorted. These projects are listed first on the row and marked by id. This is the strongest key: it needs no email and is never overridden.
email
Product users on the email's domain, through their roles, give the row its live projects. On a free-mail domain (gmail.com and the like) only the exact address counts, since the domain is shared by strangers. Emails of our own staff are kept but never matched. At most 25 projects are shown per row; a company domain can lead to thousands.
batch
The YC batch, cleaned on upload: Winter 2024 becomes w24, an empty cell null, anything unreadable stays as typed so you can fix it. It never matches anything by itself; it is the filter on the sheet and, with the email, the identity of a row when ingested.
nameor company
A fallback. When neither the ids nor the email lead to any project, and a tracked company has this name (trimmed, any case), that company's projects are listed instead, dashed and marked suggested: the founder filled in a personal email, the company runs on its own domain.

Precedence

  1. project_ids are listed first, in id order, whatever the email says.
  2. Then the projects the email leads to, in id order. A project both keys reach is shown once, with the user and role from the email.
  3. Only when both are empty does the name suggest a tracked company's projects.
  4. batch never changes the match; it filters the sheet and keys the ingest.

Editing an email or project_ids cell drops the row's cached matches, so the next open recomputes them. The tracked company owning a project is looked up live on every open.

Ingest

The Ingest button in a CSV's header writes every row into the database table silver.batches, one row per (batch, email): the row's name, the project ids it keys on or its email led to (suggestions excluded), the tracked company owning them, and the whole row as JSON. A row already there is updated, so ingesting the same file twice, or a corrected version of it, is safe. Rows without an email, or with one of our own, are skipped; the counts are shown next to the button.

Query it as bi.batches in the SQL tool, for example SELECT batch, count(*) FROM bi.batches GROUP BY 1 ORDER BY 1.