Gambling Our Future On A Story

One theme of the Trump presidency involved a continual surprise that, yes, in fact he could legally do a certain thing, with only historic precedent and informal agreements keeping such events from transpiring in the past. New pilots, meanwhile, find themselves dangerously over-fitting their models of weather, sites, regions, and people to limited experience, yielding the advice that “it’s not just the hours flown but also the years”. Our dangerous capacity for narrative fallacy — the construction of compelling stories from sparse data — knows no bounds.

In all domains our ability to develop quality theories about how they work centers on how long we have been swimming in that pool relative to the periods of relevant cycles. The same indeed holds true for technology ecosystems but their self-reinforcing nature is now collapsing cycle time so dramatically as to convert benefit to curse, giving us plenty of opportunity to see many cycles while depriving us of time to absorb each one’s import. Woe betide large governments, legacy industries, and primate brains attempting to grok the Artificial Intelligence explosion.

The notion that “good artists borrow and great artists steal”, whether attributed to Pablo Picasso or T. S. Elliot, has long existed as an acceptable yet gray reality, with charges of blatant plagiarism holding water only in the case of long strings of word-for-word (and later bit-for-bit) copying while most else remained ambiguous. And for most of human history such “stealing” has operated in a very distributed and extremely heterogeneous wetware architecture characterized by low bandwidth and high latency. In creative professions, employers and clients typically own “Works For Hire” while the developed skills and techniques remain with the creators unless otherwise stated.

The advent of the printing press certainly complicated matters by rendering copying scalable, and the Internet took that to another level by doing the same for transmission, but it wasn’t until the advent of Large Language Models, the requisite hardware, and a few related techniques that we could cram all the output of all the humans into a single “brain” that could in turn semi-autonomously begin to generate derived content. The “fair use” laws and traditions that emerged from centuries of experience informed by the limits of individual human brains or at most low-tech production conglomerates have in the last two years begun to break.

In Data Engineering we speak of “The Three V’s” — volume, velocity, and variety — where even small increases in any area will strain a human brain and increases across all three may yield a system for which no human precedent exists. As sensors have become ubiquitous, storage and networking costs have asymptotically approached zero, and (most) compute has become commoditized, the ability to consolidate and transact on a brain meltingly intense all-three-Vs collection of data in one place has upended our ability to reason about fair use, representative democracy, competitive markets, and sustainable ecosystems.

If you are the New York Times or Scarlett Johansson then you perhaps wield enough brand power, both recognition and resourcing, to fight back against the technology companies that left unchecked would gladly assimilate and profit from your creative output. If, however, you identify as, say, an individual blogger, artist, model, or similar then you inhabit a woefully vulnerable position within the present ecosystem, one subject to voracious tech behemoths who would happily and mercilessly grind you into feedstock for their AI beasts. And if, as we all ought, you have a stake in the future of our creative ecosystem, then you should be concerned that we are mortgaging our tomorrow by eating our seed corn today, a path to stagnation and instability owing to individual contributors finding their situations increasingly untenable.

This “seed corn” issue, meanwhile, extends beyond the problem of today’s creators not benefiting from their works to tomorrow’s creators simply not existing. The Software Engineering profession, as with most other crafts, has long operated under a tradition where practitioners proceed through a series of apprenticeships as they make the journey from neophyte to master. The advent of LLM-powered code generators, meanwhile, stands poised to cut off today’s junior developers at the knees, doing so by suddenly creating a treacherous valley in their potential career paths. If today’s would be junior developers cannot survive, then the three years hence mid-level developers may not exist. For want of those mid-levels, whence the subsequent seniors and principals?

Herein lies great peril when you consider that the best usage of Generative AI comes from experienced practitioners who know the right questions to ask and possess the depth to critique its output. Will we skillfully navigate the present crisis to a near future where man and machine have broadly merged and the benefits thereof prove diffuse or will we blunder headlong into one where machine intelligences accelerate away from the human masses yet also labor in thrall to a plutocratic minority? Perhaps more so than ever in human history the decisions of the next decade may cement our future as a species in perpetuity.

Whoops — this model needs more training data.
Okay — this model seems adequately trained but maybe only on this specific wing.

Discover more from All The Things

Subscribe to get the latest posts sent to your email.

Leave a Reply

Scroll to Top

Discover more from All The Things

Subscribe now to keep reading and get access to the full archive.

Continue reading