bioconductor.org moved, and the machines are gone
What bioconductor.org was running on, the incidents that forced a change, the stack that replaced it on 2026-09-28, and what one day of request logs says the site is for.
Bioconductor is many things to many people, but it is perhaps most recognizable by its website, bioconductor.org. In this first of a series of posts, I’m going to outline changes to the site and several other related pieces of infrastructure that were tied to the site. bioconductor.org is no longer served by a virtual machine in AWS. Since about 20:09 UTC on 2026-09-28 it has been served by a Cloudflare Worker reading from object storage, and BiocManager::install() kept working through the switch. I overstated a bit in the title: the old machines are still running, but they no longer carry the load, and they will be retired once the migration is complete. In this post I describe what the old site was, why it had become hard to keep running, what replaced it, and what the first days of request logs say about what bioconductor.org is for. Later posts will take the pieces one at a time.
The numbers here come from the project’s documentation site, seandavi.github.io/bioc-infrastructure, and where one is an estimate I say so. I’m adjacent to the Bioconductor core team and have done this work with their cooperation. The content of the site still belongs to Bioconductor, and decisions about it are theirs.
What the old site was
The simplest way to describe bioconductor.org is that it is a package repository with a website attached. Every BiocManager::install() reads config.yaml, PACKAGES and VIEWS from it. Mirror operators rsync the whole tree. Package maintainers read build reports under checkResults/. renv and install_version() read Archive/. The pages a person reads in a browser are the smallest share of the load.
Figure 1 shows how it was put together when I did a read-only survey of the two servers on 2026-08-03.
Two machines did everything a visitor saw. Staging was a build box. Every hour, a cron job pulled the website repository, did a full clean rebuild of a Ruby static-site generator (Nanoc, no incremental step), and rsynced the output over SSH to the other machine. About twenty other cron jobs on the same box generated landing-page JSON, badge images, build-result feeds and the search index. The results-tracking app for the single package builder ran there too, as three Python daemons started from @reboot cron with nohup. If one crashed, nothing restarted it. Its logs had been growing unbounded since about 2024, and the operating system was past end of standard support.
Master was one EC2 instance running plain Apache. Everything a visitor got was a file already sitting on its disk, put there either by staging’s hourly push or by a build machine writing check reports directly into checkResults/. The URL logic lived in a 19 KB .htaccess with hundreds of redirect and rewrite rules accumulated over about twenty years. Two things were proxied rather than served from disk: site search, to an unmaintained Solr, and the download statistics, to a separate host. CloudFront sat in front, so every cache miss landed on that one VM and its EBS volume.
Behind those two were the build machines themselves, named hosts (nebbiolo1, nebbiolo2, a few Macs) that run R CMD check across the whole repository every night and have done so for as long as the project has existed, with older names (moscato, zin, morelia, oaxaca) still in the configuration from earlier generations of hardware.
The site’s repository still carries traces of its earlier lives. A migration/ directory documents the previous move, from Plone to Nanoc, wget script and all. A file of 331 more rewrite rules has redirects for a 2002 workshop in Heidelberg among its oldest entries, though no build step refers to it. And the statistics crontab still has the Squid logs from the old Fred Hutch proxies in it, commented out rather than deleted.
A static-file server behind a CDN is a sound design for a package repository, and this system served the community reliably for a long time. What it had become by accretion was harder to run: three operating systems to patch, a build box whose daemons nobody supervised, a hardware refresh cycle for the builders, hundreds of rewrite rules nobody could safely edit, and essentially no written description of how any of it fit together. Each piece was bespoke, each lived on a particular machine, and the knowledge of how to operate it lived in a small number of people’s heads. That made it hard to migrate, because you couldn’t move a piece without first finding out what it did.
The incidents
What forced the issue was crawlers. The origin ran off an EBS volume, and crawlers exhausted its IOPS. Figure 2 shows the path.
The requests that hurt were not page views. They were bots pulling hundred-megabyte lecture videos from /help/course-materials/, where two years of materials are 711 MB of .mp4 against 2 MB of HTML (the panel on the right of Figure 2), and bots walking unique URLs that no CDN could ever have cached. That one directory was taking more than 500,000 requests a day, and when we looked on 2026-07-27 CloudFront was answering /help/course-materials/ with X-Cache: Miss, so all of it went to origin. Because one disk sat behind everything, a crawler pulling videos slowed BiocManager::install() for everyone. This came to a head over the four days up to 2026-07-27, the day of the first commit in bioc-edge, the repository for the Worker described below. During those days the site was often unresponsive or took a very long time to load, and several people noticed. The core team spent four days working with AWS staff to try to stop the bot traffic.
The stopgap was to buy more IOPS, which cost more every time and fixed nothing structural. The reasonable question was whether a bigger VM would do. I don’t think it would have, for reasons that have little to do with the disk. Pages, installs, mirrors and build reports shared one failure domain. Nobody could see the traffic: the request logs fed statistics jobs, and nobody read them request by request. The download statistics have never filtered bots; the code that would do it has been commented out for years. Most of the roughly $5,000 a month in AWS was egress, which is hard to avoid when the job is sending tarballs to the world. And nothing was written down where a newcomer could find it.
The new architecture
The decision was to move serving to object storage plus edge compute, and to replace the legacy pieces one at a time behind the live site, with every choice written up as an architecture decision record. The pattern is sometimes called a strangler migration: the new system sits in front, answers what it can, and falls back to the old one for everything else, until there is nothing left to fall back to.
As Figure 3 shows, a request now goes through Cloudflare’s firewall and edge cache, then to a Worker, a small TypeScript program that runs in Cloudflare’s data centres. The Worker looks up the path first in the latest build of the website and then in a mirror of master’s files, both stored in an R2 bucket of about 5.1 TB (1.65 million objects, including 4.66 TB of old releases). Object storage has no symlinks, so the 145 symlinks in the old docroot (packages/release pointing at 3.23, for instance) are a JSON file the Worker reads. The hundreds of .htaccess rules became a generated redirect table. A release roll is a data change, not a deploy.
Packages still come from where they always did. The Bioconductor Build System builds the source tarballs, r-universe builds most of the Windows and macOS binaries, and the core team’s propagation puts them on master. An hourly job copies master into R2 and purges exactly the URLs that changed. The website is built separately, by Astro, on every merge to the website repository, into an immutable folder named by commit. Rolling back a site change is writing an older commit id into one pointer. Every pull request gets a preview on the real worker, and the response carries a header saying which build answered it. Packages and the website meet only in storage; neither waits on the other.
The cutover
The migration ran for about two months before DNS moved, and the mirror was live and checked long before anyone depended on it. Figure 4 gives the dates.
R2 was loaded on 2026-08-03 and verified against the archive with zero differences. By 2026-08-13 the hourly sync, a weekly checksum reconcile and request logging were running with alerts. On 2026-09-28 the nameservers moved to Cloudflare at 17:30 UTC and the site flipped at about 20:09. The BiocManager::install() acceptance checks, eight of them, passed on release 3.23 and devel 3.24 after the flip. As of this writing the old DNS zone is kept, frozen, as the rollback until 2026-10-12. The next day a probe compared 13,271 paths from real traffic against master, and the differences it found (a nine-byte 404 body, package files skipping the edge cache, double-slash links, landing pages missing their Windows and macOS download links) were fixed the same day.
On cost, two numbers that are not estimates of the same thing. The AWS estate being retired is about $5,000 a month, for CloudFront, S3 and the two servers, not counting the builder hardware or anyone’s time. The new stack, at Cloudflare’s list prices applied to measured traffic, is roughly $240 to $500 a month, of which storage is about $77. That is a projection, not a bill, and until the AWS side is switched off the project pays for both. The difference is mostly egress, which R2 does not charge for.
We also read the request logs within the hour of the flip, which the old setup had never let anyone do.
In the first nineteen minutes, 75% of all requests were one crawler: 1,130 addresses on two cloud networks in Singapore, rotating three Mac Chrome user agents. It was stuck in a loop we were feeding it. The old site redirected anything under /talks to the course-materials index, we had copied that rule, and the index has 392 relative links that the crawler resolved against the URL it had asked for. Every redirect minted 392 new URLs that redirected back. No human had ever used that redirect. We removed it at 20:58, and when the crawler kept working through its queue at 3,700 requests a minute, added one firewall rule at 22:34 scoped to those two networks and that one path prefix (Figure 5).
The old site almost certainly had the same crawler, since production had the identical rule, but nobody could see it.
What bioconductor.org is for
Two days after the cutover, 2026-09-30, was the first full UTC day with clean logs. Figure 6 shows what it looked like.
5.8 million requests, 8.1 TB sent, from 215 countries, with the busiest hour at 09:00 UTC when Europe is at work and Asia’s afternoon overlaps. About 21,000 distinct addresses sent R’s own user agent that day, split by operating system in the last panel of Figure 6. Addresses are not people (one university behind NAT counts once, one laptop on three networks counts three times), but it is the first time the project has had a number like that measured directly rather than inferred from tarball downloads.
A third of the requests came from R and package clients, the first bar in Figure 6. Another third came from what our first-pass classifier calls other automation, which is mostly machines in clouds presenting browser user agents: half of that is about 740 Google Cloud addresses whose browser string carries Google’s front-end marker, and most of the rest is browser-shaped traffic from hosting networks. A fifth was browser-like, and I would not present even that as a count of humans; the top countries by distinct address in that class look more like crawlers on residential connections than people. Build reports, the pages maintainers check each morning, took 1.27 million requests from 242,000 addresses, about 22% of all traffic. Declared search and AI crawlers were 8%.
One machine downloaded 162,163 package files that day, every one of them for Bioconductor 3.19, a release two years old. That was 40% of everything R clients downloaded. It is almost certainly a CI job or a script reinstalling in a loop, and the published download statistics would have counted every one of those as a download. August and September 2025 in the published tables show downloads roughly quadrupling with no change in distinct addresses, for what is almost certainly the same reason. Those months are still inflated, and cannot be corrected, because the pipelines that produced them discarded the user agent before storing anything.
So by request count, about two thirds of the traffic is machines talking to machines, and a good part of the rest looks like automation using browser user agents. Some of that is the job: R clients, mirrors, CI systems and build reports are what a package repository exists to serve, and the AI crawlers and search engines that index package documentation are, increasingly, how people find packages. Some of it is waste that nobody could see. The distinction that matters is between traffic that serves the community’s purpose and traffic that doesn’t, and you can only draw it with request-level logs and a classifier you can argue with and improve. I think a lot of running a scientific software repository from here on will be telling these apart, and I’d rather do it from the logs than by guessing.
Where things stand
bioconductor.org is served by the new stack. Packages are still built where they always were and copied hourly from master. The download statistics are still generated by the legacy pipeline and proxied through. A registry built on r-universe, which applies one propagation gate to every package and would let the site serve packages with no dependency on the old servers, is running but not yet in the serving path. The data, experiment and workflow packages that r-universe does not build need a build system of their own, which is designed and not built. The content of the website still comes from the Bioconductor/bioconductor.org repository through a snapshot; who owns it long term is a core team decision. The old DNS zone goes away after 2026-10-12. All of this, with dates, is on the roadmap.
The posts that follow will take on some of the components in turn, such as the request logs and what a traffic classifier for a package repository should look like, the edge and how a release roll became a data change, the website as a pull-request workflow, the registry and its gate, and how the agentic workflow was set up. If you run infrastructure for a scientific software project and any of this is useful, or wrong, I’d like to hear about it: seandavi@gmail.com, or an issue on bioc-infrastructure.
Thanks to the army of volunteers and contributors who built the site and to the Bioconductor core team, in particular Lori Shepherd, Jennifer Wokaty, Hervé Pagès and Marcel Ramos, for access to the servers, patience with the questions, and the propagation and build systems that still do the hard part, and to Jeroen Ooms, whose r-universe builds most of the binaries bioconductor.org ships.