
Sci-Hub, Anna’s Archive, Library Genesis – we are all familiar with these names. They are collectively known as shadow libraries; who knows, perhaps because they only exist in the shadows. No sooner have they been identified, they are promptly shut down, only to reappear at a different location. Are they just a minor irritation for the big publishing companies, a game of whack-a-mole? I would argue that they have a more serious function than that, and are well worth studying.
Zakayo Kjellström, a trained librarian from Sweden, would appear to agree – he has a highly readable PhD on the subject (Black Open Access: shadow libraries and text piracy), and I talked to him recently about his study, for the ATG Podcast.
Of course, what the shadow libraries have in common is that they disregard intellectual property. They are collections of as much content as possible, both scholarly and otherwise, which is why publishers are so keen to have them shut down. But if we think back to the days of Napster, launched 1999, there was for a couple of years, a free-for-all for commercial music, using peer-to-peer file sharing. Almost any music track or album was available for download. Although Napster was closed down by 2001, the concept of all music being available from a convenient single source was so appealing, it is no surprise that a not dissimilar service, Spotify, launched soon after (in 2006). It wasn’t peer-to-peer, and it charges a hefty subscription, but it has become ubiquitous, simply because it is so easy to find any track at a single location. Who needs record stores any more?
Sadly, there is no equivalent for libraries. The idea of a universal library, that contains every book ever published, is a very old one, and perhaps more myth than reality. Estimates differ for how many books The Library of Alexandria held, but we remember it because it was thought to be the largest library. Fifty years ago, when prospective students searched for an institution to study, the number of books in the library was one of the criteria used, with the implication that the more books, the better the institution. Today, less attention is paid to the ownership figure, since much, perhaps most of the content is licensed.
But there is certainly a value that can be placed on comprehensiveness. A fundamental part of the systematic review for medicine is to locate all the papers that have been published on a topic – something that turns out to be more complicated than at first thought. None of the existing scholarly collections, such as Google Scholar, Web of Science, or Science Direct, is complete, and there is a significant cost in hours spent by information scientists trying to make sure they have accessed all the relevant papers.
Even the shadow libraries are not complete, but they fix this by collecting each other’s content, and most of all, they are wonderfully convenient. As Kjellstrom puts it, “On Anna’s archive, you go in, you search for the book, you press download, you wait a few seconds, and then it’s on your computer.” What you really want, says Kjellstrom, is a button in the library that says “download to my Kindle”.
Attempts to create a universal library have been seen with distrust – specifically, see the resistance to the Google Book Project, as described in Along Came Google by Marcum and Schonfeld. Yet at the same time, if you ask any researcher how much content they need access to for their research, they will tell you they need to see every published article or book, regardless of copyright restrictions. After all, that’s the point of academic research, isn’t it? How valuable is a study that only uses half the evidence?
Even if we are not researchers, however much people complain about LLMs, we are all increasingly addicted to the idea of asking LLMs, or Google, or both, any question in the universe, and expecting an accurate answer. For the LLMs to do this, they require a complete collection of the world’s content, with no restriction on copyright.
So the shadow libraries win both for universality and for convenience. Of course, they don’t solve any problem of IP in the long-term – although they claim to be democratizing knowledge. Nonetheless, if an article or book is available on shadow libraries, it will be more read (a 2021 article in Scientometrics found papers on Sci-Hub were cited 1.72 times more than articles not available on the site). In contrast, legal efforts to make academic content more available over the last 20 years have not transformed scholarly publishing. Diamond Open Access still represents a tiny proportion of the entirety, although Kjellstrom ends on an optimistic note: “I genuinely think the continuing of small-scale initiatives is really the way forward, and creating communities around them, so you don’t feel so alone in your own institution or library or publishing house”. This may or may not turn out to be scalable, but in the meantime, it is paramount to examine the present-day state of affairs, as Kjellstrom has done.

Leave a Reply