Research software & data infrastructure · Biological data provenance
Lamin
An open source lakehouse that tries to do for biological data what git did for code: record where every dataset came from and what produced it. The code is public and auditable. The company behind it is almost entirely not.
A Slack channel and a mountain
In the autumn of 2023 a PhD student named Jérémie Kalfon was stuck. He had set out to train a foundation model on single cell RNA sequencing data, largely alone, and had hit the thing that stops most such projects. Not the model. The data.
“Managing a dozen datasets is already painful for most computational biologists. I needed to handle thousands,” he later wrote. Every study used different gene identifiers, different tissue labels, different annotations.
He found Alex Wolf in the Chan Zuckerberg Initiative Slack channel for CELLxGENE, circling the same question. Wolf had a year-old company working on exactly that.
Two and a half years later the answer was in print. Kalfon trained scPRINT on 50 million cells and published it in Nature Communications in April 2025. The paper states: “Parallel to this work, we worked with Lamin.ai to develop a dataloader for large cell atlases.” On his own blog Kalfon put it less formally: “Without LaminDB I would have spent months on this. With it, it took weeks.”
What the pile looks like
Biology does not produce one big table. It produces enormous numbers of small, awkward, partially overlapping ones: a sparse expression matrix annotated down both rows and columns, a folder of microscope images, a sequence file, each labelled against a vocabulary the next lab uses slightly differently.
Standard infrastructure was not built for this. A commissioned explainer on Lamin's blog put the trilemma cleanly: “Warehouses are too rigid, data lakes can’t be queried, and tabular lakehouses don’t understand the formats.” Iceberg and Delta are excellent at tables and blind to a zarr store.
LaminDB aims at the missing middle. Metadata lives in SQLite or Postgres. The data stays where it is, in parquet or zarr or AnnData, on disk or S3. What the system adds is bookkeeping: which run produced which artifact, from which inputs, under which code version, against which ontology term. The framing has moved with the market. The repository now opens on agents: “Untraceable results cannot be trusted, especially when non-verifiable tasks are delegated to agents.”
Two founders and one prior act
Alex Wolf holds two doctorates: electrical engineering from Erlangen-Nuremberg in 2014, after modelling solar cell chemistry at Bosch Research, and computational physics from LMU Munich in 2015, which won the department prize for best thesis. He then spent three years in Fabian Theis's machine learning group at Helmholtz Munich.
That is where he built the thing he is known for. Scanpy, started in 2016, became the dominant Python toolkit for single cell analysis. Its 2018 Genome Biology paper has 8,582 citations in Europe PMC; his trajectory method PAGA has 1,546; he co-authored scVelo, at 2,818, and co-created anndata, whose h5ad format became the way the field moves data. Announcing a 2024 Frontiers of Science Award for the Scanpy paper, Helmholtz Munich noted that “Since its release in 2018, Scanpy has been widely adopted by the research community, contributing to numerous clinical and other studies.”
In 2018 he left academia for Cellarity, the Flagship Pioneering drug discovery company, as the first computational hire and third employee overall. He stayed until early 2022, rising to Head of Scientific Computing while the company went from three people to roughly a hundred.
His co-founder Sunny Sun, formally Xiaoji Sun, came out of Jef Boeke's lab at NYU, on synthetic yeast chromosomes and LINE-1 retrotransposons. She is second author on the 2018 Nature paper showing that karyotype engineering by chromosome fusion produces reproductive isolation in yeast. Wolf's CV records that she joined his team at Cellarity in April 2019; she later ran computational biology there.
Their account of why they left is short. “While running ML and Comp Bio teams, we learned that the biggest obstacle to better AI in R&D was closing the feedback loop with the wetlab,” the about page says.
Four years of commits
The repository was created on April 15, 2022. Y Combinator took the company into its Summer 2022 batch. A problem statement went up that July, opening with a line the company has kept for four years: “The complexity of modern R&D data often blocks realizing the scientific progress it promises.”
What followed was volume. By September 23, 2026 the main repository carried 5,373 commits and 283 tagged releases, the org 47 repositories, the package version 2.10.0. There is an R client, a Nextflow plugin, and an ontology package, bionty, downloaded more often than lamindb itself. The 2026 work reads like a land grab on public data: a queryable mirror of the Arc Institute Virtual Cell Atlas, 2.5 billion expression profiles that Arc ships as 460,000 files across 41 terabytes, and native support for SpatialData and Vitessce.
What is proven, and what is still claimed
| Evidence | What the record shows | Source type |
|---|---|---|
| Open source repository | laminlabs/lamindb created Apr 15, 2022 under Apache 2.0. At Sep 23, 2026: 5,373 commits, 283 releases, 23 contributors, 368 stars, 68 forks, last push the same day. | Public record |
| Package downloads | lamindb drew 28,510 PyPI downloads in the 30 days to Sep 23, 2026, bionty 32,519. For scale, Wolf's earlier scanpy drew 617,692 and anndata 1,108,517. | Public record |
| Founder publication record | Europe PMC, Sep 23, 2026: SCANPY (2018, first author) 8,582 citations; PAGA (2019, first author) 1,546; scVelo (2020) 2,818; scverse (2023) 361, with Lamin Labs listed as an affiliation. | Public record |
| Independent use in the literature | Kalfon et al., scPRINT, Nature Communications 16:3607 (2025), credits joint dataloader work with Lamin.ai. Europe PMC returns 3 records mentioning LaminDB, all from outside groups. | Independent |
| Named users | The README lists Pfizer, Altos Labs and Ensocell Therapeutics; scverse, DZNE and Helmholtz Munich; and a swarm learning network naming Harvard, MIT, Stanford, ETH Zürich and Mount Sinai. None has confirmed a contract or spend. | Company-stated |
| Scale and compliance claims | The README gives order of magnitude figures for pharma deployments and cites GxP, 21 CFR Part 11 and EU Annex 11 as drivers. No customer confirmation, no published validation or audit certificate, no FDA record of any kind. | Not found |
| Paper describing LaminDB itself | None. Europe PMC and PubMed searches on Sep 23, 2026 return no peer reviewed methods paper from the Lamin team. | Not found |
| Revenue | An AI generated aggregator states $3M revenue while also calling the company pre-seed. PitchBook's revenue field is blank. No filing or company statement supports a figure. | Unsupported |
| Filings and federal funding | SEC EDGAR full text search returns zero hits for Lamin Labs, LaminDB or lamin.ai across all form types. No NIH or NSF award found. | Public record |
Read plainly: the technical record is strong and the commercial record is thin. Four years of daily commits under a permissive licence, a founder cited more than eight thousand times for his last tool, and an outside group that used this one to train a published foundation model is real evidence the thing works. What is missing is anyone outside the company saying what they pay or how much data they run through it. The company names Pfizer. Pfizer has not named Lamin.
What to watch
- A named, confirmed enterprise deployment. The customer list is the largest claim and the least corroborated.
- Contributor concentration. Twenty-three contributors on a project this size means a paid team writes nearly all the code.
- Whether the agent framing converts. Downloads, not positioning, will show it.
- A first Form D, or any priced round visible in EDGAR. Four years in, there is nothing on file.
- A methods paper. In this field a citable paper is how adoption compounds. Scanpy's did exactly that.
In their words
“While running ML and Comp Bio teams, we learned that the biggest obstacle to better AI in R&D was closing the feedback loop with the wetlab.”
Lamin about page, 2026 · Company
“Untraceable results cannot be trusted, especially when non-verifiable tasks are delegated to agents.”
lamindb repository README, 2026 · Company
“Managing a dozen datasets is already painful for most computational biologists. I needed to handle thousands.”
Jérémie Kalfon, author of scPRINT, personal blog, 2026 · User
“These days, when I start a computational biology project, I set up a git repo and a LaminDB instance. In that order, roughly.”
Jérémie Kalfon, personal blog, 2026 · User
“Parallel to this work, we worked with Lamin.ai to develop a dataloader for large cell atlases.”
Kalfon, Samaran, Peyré and Cantini, Nature Communications, 2025 · Peer reviewed
“Warehouses are too rigid, data lakes can’t be queried, and tabular lakehouses don’t understand the formats.”
Jesse Johnson, commissioned post on the Lamin blog, Mar 2026 · Company blog
“It’ll be interesting to observe which systems of record agents will ultimately prefer, and we’ll keep optimizing for that at Lamin.”
Alex Wolf, Symbolic memory for biological R&D, Feb 2026 · Company
Related companies
Sources
- Public recordGitHub API record for laminlabs and related repositories
- Public recordPyPI download statistics
- Public recordEurope PMC citation records for F. Alexander Wolf and Xiaoji Sun
- Public recordSEC EDGAR search, and Lamin Labs GmbH imprint HRB 279584
- Peer reviewedscPRINT: pre-training on 50 million cells
- Peer reviewedThe scverse project
- User accountHow I managed thousands of datasets to build scPRINT
- InstitutionHelmholtz Munich Team Wins Frontiers of Science Award
- IndependentLamin company and funding profiles
- Companylamindb README, user list and scale table
- CompanyLamin about page, current and 2023 archive
- CompanyLamin blog: Key problems, Symbolic memory, Sparse measurements
- CompanyAlex Wolf curriculum vitae
- Public recordThe Y Combinator Standard Deal
Profile researched and written by Healthcare Discovery. Last updated September 29, 2026.
