What Payments Taught Me About Genomic Data
Science has a limited supply of doubt, and genomics spends most of it on the file.
The checkout cart
My family started a water delivery business and needed a way to take money for the water, so I built the checkout.
The first version of the cart treated every card or bank failure the same way, which was to say “it did not work”, go call the customer and profusely apologize. We quickly discovered that this was not a sustainable error-handling method. The failures, it turns out, actually come to our system with a reason attached. “Insufficient funds” is not an “expired card” is not a “network timeout”. One of those will succeed if I try it again in a minute, and two of them never will. (Well maybe two of them could succeed.)
This was news to me. My day “job” at the time was a PhD, which is to say screening twenty thousand strains across a handful of conditions, and nothing in six years of genomics had ever arrived with a reason attached. A result that looked wrong did not say why it looked wrong. It sat there, and my method, which I believed in without ever quite saying so out loud, was to dig through the dataset with enough fervor that I could eventually get the data to confess, which in practice meant plotting it again, from a different angle, on a different axis, in a different color. I made a lot of plots that year.
The payment processor just told me, quite calmly, in two words. All it asked of me was to teach my cart to speak the right language, and then to get on with my day.
Doubt is a budget
The customer had paid or the customer had not paid, and the system knew which and said so, and then I wound up taking this for granted within about a week. When the connection dropped halfway through a charge, which it did, the office being a kitchen table, I could press the button a second time, and the system would recognize the same order and tell me it had already been paid, which meant that nobody paid for their water twice. In the lab, if I had a dropped connection or if I forgot, I’d just copy the file again, and then I’d have two, one being called File (1).
And when a card failed, I was told why. A bank I had never heard of decided that a card was empty, and said so to a network, and the network said so to a processor, and the processor said so to me, at a kitchen table, in the same words the bank had used, and at no point along the way did anybody have to remember it, or be asked. I did not build any of it. It was there when I arrived, like the plumbing.
That left my doubt for the things that deserved it. The decline code has nothing to say about whether anyone actually wants water delivered to their door, or whether I have priced a five-gallon jug correctly, or whether my family ought to be in the water business at all, which is still not unanimous, by the way. Nobody emits a code for that. Those were the questions worth being unsure about, and for once I had my whole attention to be unsure with.
Doubt is a budget. In genomics I had been spending most of mine on whether a file had finished copying or when the server might run out of space. And a file that has finished copying is the easy case, because somebody can at least say what finished looks like. Nobody has ever told me what a good sequencing run looks like. There are terabytes of it, far more than anyone could look at, so what you wind up looking at is a plot, and then another plot, for weeks, each one written in R with slightly more dread than the one before, since the code that draws it is also the code that will tell me whether there is anything here worth studying at all. At the end of it all, nobody says pass or fail. Am I just playing with noise?
All else equal
The irritating part is that I came to genomics to get away from exactly that question. I came from economics, where ceteris paribus gets a small laugh in every lecture, the kind a room gives a joke it has heard from every professor it has ever had, and I laughed along for about six years of coursework without knowing what was funny. Then I took a job as a natural resource policy analyst, where my work was to put an economic cost on endangered species policy, which in practice meant a small fish, a river without much water in it, and a spreadsheet in which the fish was either listed or it was not, and the difficulty is that there is one New Mexico, and we listed the fish in it, and nobody has yet produced a second New Mexico where we did not, so that I could compare the two. So I built a model, and the model was handwaving with a citation attached, and it went into a report, and the report went to a meeting I was not at, and somebody, presumably, made a decision with it. I got the joke eventually.
I quit, and very nearly walked straight into medicine, having finished most of the prerequisites, organic chemistry included, before a professor asked me what exactly I thought a physician does, which is to get one patient, once, who cannot be run again with one variable changed and who has to be treated anyway, usually by the end of the appointment. She mercifully sent me to genetics, where I did not know bioinformatics existed until I was already doing it, and what I found there was bacteria, which are the closest thing biology has to bare metal data.
In a pooled CRISPRi screen, twenty thousand strains sit in the same well, in the same medium, at the same temperature, in the same dose of the same drug. They differ in one respect, which is the gene that is knocked down in each of them. Ceteris paribus is the design of the experiment, and the answer comes back in a day, because the phenotype is built in. Grow or die. Here, finally, was a place where I could hold the world still and turn one knob, and then turn a different one in the next well, because there are ninety-six wells on a plate and nothing stops me from running another plate tomorrow.
Economics gets one New Mexico. I get as many variables as I can pipette.
Down the hall
Designing the experiment is a good reason to trust the biology. It is not a reason to trust the file. The data comes off an instrument, gets copied to a server, gets copied again onto the laptop of a rotation student who is in the lab for four weeks and who names his directories after the day of the week he made them, and from that point on the only record of what happened to it is whatever he remembers, which is admittedly quite a lot, including that one of the runs had to be redone and he is fairly sure it was the second one. Was the transfer complete? Did every sample make it? Is the copy on the cluster still the same as the copy that went to the collaborator? No system will answer any of that. There is only asking somebody who was definitely paying attention at the time.
Well, it turns out that rotation student was me! The directories really were named after the days of the week, and I only stopped because I ran out of week.
Partway through my PhD I built a server and started keeping other people’s data on it, first for my own lab, then the labs next door, then PIs I had never worked with, then visiting scientists. I called it a server, and after a while everyone else called it a server, and it was an Ubuntu workstation that I talked my PI into letting me build and then stick in a room down the hall. People kept coming back to it anyways, because sometimes I kept track of what was on it and where each piece had come from, or I could press up enough times in my bash history to recreate it, and keeping track was nobody else’s job in the building.
My backup strategy was that I did not delete things.
I knew, with some pride, where every file on that machine had come from and who else had a copy, and it did not once occur to me to write down why a run had come back empty. I had somewhere to put the data and nowhere to put the reason. The record of why a run came back empty was that I had been standing there when it did.
Tellus
I have built some version of this three times now, and it took me that many to work out what I was building. barcoder is a script. There are a few files, I run it, it finishes, and the record of what happened is that I was sitting there when it did, which had always been good enough for me and was good enough for barcoder. crispr_experiment was for other people, and a person designing a library against a whole genome is not watching a script for a few seconds. They start it, go and do something else, and come back to a result they did not see being made, and it was only then that I began to care whether the thing could say what it had done. Every version asked the person using it to write a little less code.
Tellus, which Nurture has just opened in private beta, is the one where the scientist does not have to write the plot. She opens her data, all the terabytes of it, and points at the part she wants to see, and it is there, and she points at another part, and that is there too, and at no stage does she cycle through a folder of plotting scripts trying to remember which of them she changed last. I am aware that pointing and clicking has never made anybody’s pulse race. But I used to spend weeks finding out whether I was playing with noise, and it turns out that can be found out before lunch. And when she hands the data to the next person, the next person can point at it as well, which is an improvement on asking a graduated PhD student to check his laptop, in hopes that his bash history isn’t completely obliterated.
My colleagues will tell you that Tellus is how the epigenome becomes computable, and they are right. I am a bacteriologist, and what I see when I look at it is the part underneath, which is a system that can say where the data is. That part does not care what kind of sequencing went into it.
Nature withholding the answer is the only part of this that was ever supposed to be hard. Everything else can be known, if somebody builds the thing that knows it. I would like to spend my doubt on nature.
Nurture is building the infrastructure for epigenomics: tools to read it, software to understand it, and systems to remodel it. Tellus is in beta. Request access if you want to run your own data through it.