arXiv

The overlay bias

Vasudevan Mukunth

Dec 14, 2020 — 4 min read

I’m not very fond of some highly popular pieces of writing (I won’t name them because I’m nervous about backlash from authors and/or their supporters) because a part of their popularity is undeniably rooted in technological ‘solutions’ that asymmetrically promote work published in the solution’s country of origin.

My favourite example is Pocket, the app that allows users to save copies of articles to read later, offline if required. Not long ago, Pocket introduced an extension for the Google Chrome browser (which counts hundreds of millions of users) such that every time you opened a new tab, it would show you three articles lots of other Pocket users have read and liked. It’s fairly brainless, ergo presumably non-malicious, and you’d expect the results to be distributed equally from among magazines, journals, etc. published around the world.

However, nine times out of ten – but often more – I’d find articles by NYT, The Atlantic, The Baffler, etc. there. I was reluctant to blame Pocket at first, considering their algorithm seemed too simple, but then I realised Pocket was just the last in a long line of other apps and algorithms that simply amplified existing biases.

Before Pocket, for example, there might have been Twitter, Facebook or some other platform that allowed stories from some domains (nytimes.com, thebaffler.com, etc.) to persist for longer on users’ feeds because they were more easily perceived to be legitimate than articles from other sources, say, a Venezuelan newspaper, a Kenyan blog, a Pakistani magazine or a Vietnamese journal. Or there might have been Nuzzle, which auto-compiles a digest of articles that others your friends on the social media have shared most – likely unmindful of the fact that people quite often share headlines, or domains they’d like to be known to be reading, instead of the articles themselves.

This is a social magnification like the biological magnification in nature, whereby toxic substances pile up in greater quantities in the gizzards of animals higher up in the food chain. Here, perceptions of legitimacy and quality accumulate in greater quantities in the feeds and timelines of people who consume, or even glance through, the most information. And this way, a general consciousness of what’s considered desirable erects itself without anything drastic, with just the more fleeting and mindless actions of millions of people, into a giant wheel of information distribution that constantly feeds itself its own momentum.

As the wheel turns, and The Atlantic publishes an article, it doesn’t just publish a good article that draws hundreds of thousands of readers. It also rides a wheel set in motion by American readers, American companies, American developers, American interests and American dollars, with a dollop of historical imperialism, that quietly but surely brings the world a good article plus a good-natured reminder that The Atlantic is good and that readers needn’t go looking for anything else because The Atlantic has them covered.

As I wondered in 2017, and still do: “Will my peers in India have been farther along in their careers had there been an equally influential Indian for-publishers tech stack?” Then again, how much is one more amplifier, Pocket or anything else, going to change?

I went into this tirade because of this Twitter thread, which describes a similar issue with arXiv – the popular preprint repo for physical sciences, computer science and applied mathematics papers (don’t @ me to quibble over arXiv’s actual remit). As the tweeter Jia-Bin Huang writes, the manuscripts that were uploaded last – i.e. most recently – to arXiv are displayed on top of the output stack, and what’s displayed on top of the stack gets more citations and readership.

This is a very simple algorithm, quite like Pocket’s algorithm, but in both cases they’re algorithms overlaid on existing bias-amplifying architectures. In a sense, they’re akin to the people who might stand by and watch a lynching, neither egging the perpetrators on nor stopping them. If the metaphor is brutal, remember that the effects on any publication or scientist that can’t infiltrate or ‘hack’ social biases are brutal as well. While their contents and their ideas might deserve international readership, these publications and scientists will need to spend more – energy, resources, effort – to grab international attention again and again.

The example Jia-Bin Huang cites is of scientists in Asia, who – unlike their American counterparts – can’t upload a paper on arXiv just before the deadline so that their papers sit on top of the stack because 2 pm in New York is 3 am in Taipei.

As some replies to the thread indicated, the people maintaining arXiv can easily solve the problem by waiting for the deadline to pass, then randomising the order of papers displayed in its email blast – but as Jia-Bin Huang notes, doing that would mean negating the just-in-time advantage that arXiv’s American users enjoy. So here we are.

It isn’t hard to see how we can extend the same suggestion to the world’s Pockets and Nuzzles. Pick your millions of users’ thousand most-read articles, mix up their order – even weigh down popular American publishers if necessary – and finally advertise the first ten items from this list. But ultimately, until technological solutions actively negate the biases they overlie, Pocket will lie on the same spectrum as the tools that produce the biases. I admit fact-checking in this paradigm could be labour-intensive, as could relevance-checking vis-à-vis arXiv, but I also think the latter would be better problems to solve.

Rapid rotation explains unusual stability of C2 anion

Tamil Nadu's lukewarm heatwave policy

The nanoscopes setting the stage for AI in biology

A view of WordPress

Read more

Rapid rotation explains unusual stability of C2 anion

Tamil Nadu's lukewarm heatwave policy

The nanoscopes setting the stage for AI in biology

A view of WordPress