Why Durable Software Takes Years to Find Its Shape
Picture a researcher staring at a library catalog. A small browser button offers to save the book, and one click produces a title, author, publisher, date, and link. The action lasts a second; the decisions underneath it—what counts as a book, which date matters, how to handle three authors, what happens when the website changes—are the real product.
Why does a tool that saves one citation require years of design? Because research software sits between messy human work and strict computer storage. It must understand enough of both worlds to make a quick action trustworthy.
Zotero is a useful case study. It grew at George Mason University’s Roy Rosenzweig Center for History and New Media, and its first public beta arrived on October 5, 2006. The release was visible, but the harder work happened earlier: conversations, rough experiments, discarded assumptions, and arguments about what scholars actually needed to keep. (zotero.org)
A citation is a data model, not a string
Zotero is a reference manager, meaning software for organizing sources and producing citations. That description sounds tidy until you look at the information a research source contains. Metadata is descriptive information about an object: its title, creators, publication date, journal, edition, pages, identifiers, and more. A schema is the expected shape of that information.
A useful mental model might look like this:
itemType: journalArticle
title: A history of software repair
creators:
- creatorType: author
firstName: Mina
lastName: Lee
date: 2026
identifier: example-doi
The difficult decisions hide in the exceptions. An author might be an organization. A book may have editors but no listed author. A web page may display a publication date that differs from the date stored in its metadata. A translator may return a plausible title while quietly losing the language or edition.
Those are not cleanup tasks to postpone until after launch. They define the product. Once thousands of people have saved records, changing the meaning of a field becomes a migration problem. A migration is a controlled conversion that moves old stored data into a new format without damaging it.
Prototype to learn, not to pretend
A prototype is a rough working version built to answer questions. In browser-based software, that might mean a bookmarklet—a saved browser bookmark that runs a small JavaScript program—or a script that captures a page and sends its contents to a server.
That first experiment can prove something valuable: people want to capture research while they browse. It can also expose the gaps. A title-and-link collector is not yet a research library. It does not tell you how to represent creators, preserve notes, attach a PDF, recover from a failed network request, or keep a user’s data safe.
The slow formation of durable software happens when a team treats those gaps as discoveries rather than inconveniences. The prototype remains useful, but the team starts separating experiments from promises. “Can we capture this page?” is an experiment. “Will this record still open and cite correctly ten years from now?” is a promise.
Stable boundaries let the world change
Durable systems usually divide their work into parts with clear boundaries. For a research collector, the flow might look like this:
web page
↓
browser connector
↓
site translator → normalized item
↓
local database
├─ optional synchronization
└─ citation output
A connector links the browser to the research application. A translator is a small adapter that knows how to read one website or data format. It turns inconsistent page content into a normalized item, meaning a record arranged according to the application’s shared schema. Zotero’s documentation describes hundreds of translators, including web, import, export, and identifier-search translators. (zotero.org)
That separation matters. If a library changes its page layout, the translator may need repair, but the database and citation engine should not. If a new citation style appears, the stored research records should remain untouched. If online synchronization fails, a local database should still hold the user’s library; Zotero describes its local storage and optional syncing in those terms. (github.com)
An interface is a small agreement about what a component must do. A translator interface could be no more complicated than this:
interface Translator {
detect(page): itemType | null
extract(page): RawRecord[]
}
The agreement creates room for change. New translators can be added without rewriting the application. Tests can feed saved pages into detect and extract. A failing site becomes a contained repair instead of a crisis spread across the whole codebase.
AI speeds implementation, not agreement
An artificial intelligence coding assistant can draft a translator from an HTML sample, produce test cases, explain unfamiliar code, and write a first migration script. That is useful work. It does not settle whether an organization belongs in the author field, whether two dates should be preserved, or which data must never be overwritten.
Those decisions belong to the domain model: the team’s shared description of the objects, rules, and relationships the software represents. An invariant is a rule that must remain true, such as “every saved item has an item type” or “a failed sync never deletes the only local copy.” A test fixture is a saved example used to check that behavior again after the code changes.
This is where AI becomes more valuable after the slow thinking has happened. Once the team knows its invariants, an assistant can generate dozens of awkward fixtures: missing authors, malformed dates, duplicate identifiers, pages with several records, and non-Latin names. It can help widen the test net. It cannot replace the judgment that decides what the net is meant to protect.
Durability includes the people
Software survives through more than architecture. It needs maintainers, documentation, contribution paths, funding, and a license that lets others understand and extend it. Open-source software means the code can be inspected, modified, and redistributed under its license. Zotero’s source is available under the AGPLv3, an open-source license designed to preserve those freedoms even when software is offered over a network, and the project operates through the nonprofit Corporation for Digital Scholarship. (zotero.org)
The work continues long after the original design period. As of October 7, 2026, Zotero’s current maintenance release was 10.0.6, with fixes touching full-text search, PDF rendering, browser-platform security, and the local API. The project also announced a move toward releases roughly every six to ten weeks in 2026. Fast updates and slow formation are not opposites: once the foundations are clear, frequent change becomes safer because it has stable places to land.
The slow part of software development is not typing code. It is discovering what must remain true when websites, browsers, teams, storage systems, and user expectations keep moving. Durable software grows from those discoveries, then turns them into data models, interfaces, tests, and stewardship practices.
That is why the best first version may take years to find its shape. The goal is not software that never changes. It is software whose changes do not erase the trust built by everything that came before.
Comments (0)
No comments yet. Be the first to respond!
Leave a Comment
Your comment will be visible after review.