文章

特德姜:人工智能不会创造真正的艺术

转一篇我非常喜欢的科幻作家 特德姜 关于人工智能艺术创作的看法。

文章使用 Claude sonnet 翻译。

原文:纽约客文章: Why ai isn’t going to make art


特德姜:人工智能不会创造真正的艺术

1953 年,著名作家 Roald Dahl 创作了一篇名为《伟大的自动写作机》的短篇小说。故事讲述了一位电气工程师的奇思妙想,这位工程师暗地里渴望成为一名作家。在他完成了世界上运算速度最快的计算机后,他突发奇想,认为 “ 英语语法的规则之严谨,几乎可以用数学来描述 “。于是,他发明了一台小说创作机器,这台机器能在短短 30 秒内创作出一篇 5000 字的短篇小说。若要创作一部长篇小说,则需要 15 分钟,而操作者只需要像驾驶汽车或弹奏管风琴一样,通过操纵手柄和踏板来调节作品中幽默与悲情的程度。这台机器创作的小说大受欢迎,仅仅一年之内,英语文学界出版的小说中,就有一半出自这位工程师的发明。

艺术真的能像达尔小说中描述的那样,仅仅通过按个按钮就创造出来吗?目前,像 ChatGPT 这样的大型语言模型生成的小说质量还很差,但我们可以想象这些程序在未来可能会有显著改进。那么,它们最终能达到怎样的水平呢?它们能否在写小说、绘画或制作电影方面超越人类,就像计算器在加减运算上胜过人类一样?

艺术一直是个难以定义的概念,区分好的艺术和坏的艺术更是困难。

不过,我可以给出一个概括性的定义:

如果 AI 要根据你的提示生成一个一万字的故事,它就必须填补你没有做出的那些所有选择。AI 可以通过不同方式来实现这一点。一种方法是取互联网上现有文本的平均值,这相当于采用了最不出彩的选择,这也就解释了为什么 AI 生成的文本常常显得平淡无奇。另一种方法是让程序模仿特定作家的风格,但这往往会产生高度雷同的作品。无论哪种方式,都难以创造出真正有趣的艺术作品。

我认为,这个原理同样适用于视觉艺术,尽管画家在创作过程中做出的选择更难量化。

真正的绘画作品中包含了艺术家无数次的决策。相比之下,使用 DALL-E 这样的文本生成图像应用的人,只需输入一个简单的提示,比如 “ 一个穿盔甲的骑士与一条喷火的龙战斗 “,然后让程序完成剩下的工作。(即便是 DALL-E 的最新版本,也只能接受最多四千个字符的提示——虽然已经有几百个词,但仍然不足以描述场景的每一个细节。)最终生成的图像中,大部分细节都是从网上类似的绘画中借鉴而来。虽然生成的图像可能看起来精美绝伦,但输入提示的人实际上并不能将其创作的功劳归于自己。

一些评论家认为,图像生成将对视觉文化产生如同摄影术诞生时那样深远的影响。这种观点表面上看起来有道理,但我们需要更深入地思考摄影与生成式 AI 之间的相似性。回想摄影技术刚刚出现的时候,人们可能并不认为它是一种艺术媒介,因为看起来似乎不需要做太多选择;你只需设置好相机,按下快门就行了。然而,随着时间推移,人们逐渐意识到相机的潜力是无穷的,

我们可以设想一个更高级的文本到图像生成器,它允许用户在多次交互中输入数万字的指令,从而实现对生成图像的精细控制;这有点类似于一个纯文本界面的 Photoshop。使用这样的程序的人,我认为仍然可以被称为艺术家。

电影导演 Bennett Miller 就曾使用 DALL-E 2 创作出一系列引人注目的图像,并在 Gagosian 画廊展出。他通过精心设计的文本提示,并反复指示 DALL-E 修改和调整生成的图像,最终从超过十万张生成的图像中筛选出二十张作品。然而,Miller 表示在 DALL-E 的后续版本中无法复制这样的成果。我猜测,这可能是因为 Miller 在用一种 DALL-E 并非设计用途的方式使用它;就好比他把 Microsoft Paint 改造成了 Photoshop,但一旦 Paint 更新,他的改造就失效了。OpenAI 可能并不打算开发一个需要用户花费数月时间才能创作出一幅图像的产品,因为这对大多数用户来说并不实用。相反,他们希望提供一个能快速、轻松生成图像的工具。

更难想象的是一个能在多次交互中帮助你写出一部优秀小说的程序。这样的程序可能需要你输入十万字的提示,然后它才能生成另外十万字组成你心目中的小说。这种程序会是什么样子,目前还很难说清。理论上,如果这样的程序存在,使用者或许可以被称为作者。但是,我不认为像 OpenAI 这样的公司会去开发需要用户付出与从头写小说一样多努力的 ChatGPT 版本。生成式 AI 的吸引力在于它能产生远超输入的内容,但这恰恰也是它们难以成为艺术家有效工具的原因。

这些公司宣称生成式人工智能可以释放创造力,本质上是在说艺术可以只有灵感而没有汗水。但灵感和汗水并不是能轻易分开的。我不是说艺术必须伴随枯燥乏味,而是指出艺术在每一个层面都需要做出取舍;在实现过程中做出的无数微小决定,与在构思阶段做出的少数宏观决定,对最后的作品同样重要。把“宏观”自动理解成“更重要”,在艺术创作中是一种误解。真正的艺术性,就存在于宏观与微观之间的相互作用里。

我怀疑,认为灵感胜过一切的想法,往往来自于对某种艺术媒介缺乏深入了解。即使是创作娱乐作品而非高雅艺术,这一点也同样适用。人们常常低估了创造娱乐作品所需的努力;一部惊悚小说虽然可能无法达到卡夫卡所说的理想书籍——” 劈开我们内心冰封的大海的斧头 “——的高度,但它仍然可以像瑞士手表一样精心打造。一部优秀的惊悚小说不仅仅依靠其前提或情节。我怀疑,如果你用语义相同但表达不同的句子替换小说中的每一句话,最终的作品很难保持同样的娱乐效果。这说明,小说中的每一个句子——以及它们代表的微小选择——都在决定整部作品的效果。

许多小说家都遇到过这样的情况:有人信心满满地找上门来,声称自己有一个绝妙的小说创意,愿意与作家分享,条件是平分未来的收益。这种人无意中暴露了他们的一个观点:他们认为写作中最麻烦的部分是组织语言,而没有意识到遣词造句恰恰是散文叙事的核心所在。生成式 AI 之所以受到某些人的青睐,正是因为它迎合了这样一种想法:人们可以在某种艺术媒介中表达自己,而不必真正掌握该媒介的创作技巧。

然而,传统小说、绘画和电影的创作者之所以被这些艺术形式吸引,恰恰是因为他们看到了每种媒介独特的表现力和潜力。正是他们对充分发掘和利用这些潜力的热情,使得他们的作品无论是作为娱乐还是艺术,都能给人以满足感。

当然,大多数写作,无论是文章、报告还是电子邮件,通常不需要作者做出成千上万的选择。那么,在这种情况下,使用自动化工具来完成写作任务有什么问题吗?让我提出一个观点:任何值得读者关注的写作,都是作者付出努力的结果。虽然写作过程中的努力并不能保证最终作品一定值得一读,但没有努力就不可能产生有价值的作品。我们阅读个人电子邮件和商业报告时所付出的注意力类型可能不同,但在这两种情况下,只有当作者投入了思考,我们的注意力才是有价值的。

最近,谷歌在巴黎奥运会期间播出了一则广告,宣传其与 OpenAI 的 GPT-4 竞争的 Gemini。广告展示了一位父亲使用 Gemini 为女儿撰写一封给奥运运动员的粉丝信。然而,这则广告引发了广泛的反对声,以至于谷歌不得不将其撤回。一位媒体教授甚至称之为 “ 我见过的最令人不安的广告之一 “。值得注意的是,尽管这封信并不涉及高深的艺术创造,人们仍然如此强烈地反对。毕竟,没人期望一个孩子给运动员的粉丝信会有多么与众不同。如果这个小女孩自己写信,很可能与其他无数封信没什么区别。但是,一个孩子的粉丝信的意义——无论对写信的孩子还是收信的运动员——在于它的真挚,而不是它的措辞有多么优美。

我们中的许多人都寄过商店购买的贺卡,也知道收信人会意识到卡片上的文字不是我们自己写的。我们不会用自己的笔迹抄写现成贺卡上的文字,因为那会让人感觉不诚实。程序员 Simon Willison 将大型语言模型的训练比喻为 “ 版权数据的洗钱 “,这是一个很有启发性的比喻:

有人认为,大型语言模型并非在 “ 洗白 “ 它们所训练的文本,而是在学习它们,就像人类作家从阅读中学习一样。但这种比较是不恰当的。大型语言模型不是作家,甚至不是语言的真正使用者。语言本质上是一种交流系统,需要有交流的意图。你手机的自动完成功能可能会提供好的或坏的建议,但它并不是在试图与你或你的通讯对象交流。ChatGPT 能生成连贯的句子,这可能让我们误以为它以某种方式理解语言,但实际上它并没有任何交流的意图。

让 ChatGPT 说出 “ 我很高兴见到你 “ 这样的话很容易,但我们可以肯定的是,ChatGPT 并不真的感到高兴。一条狗或一个还不会说话的孩子可以表达它们见到你很高兴,尽管它们不能用语言表达。

因为语言对我们来说如此自然,我们很容易忘记语言是建立在主观感受和表达欲望之上的。当大型语言模型生成连贯的句子时,我们很容易将这些人类特有的体验投射到它身上,但这实际上是一种错觉。这就像蝴蝶翅膀上的眼状斑点能欺骗鸟类,让它们误以为面对的是一个有大眼睛的捕食者。在某些情况下,这种伪装确实有效;鸟类不太可能攻击有这种斑点的蝴蝶。但蝴蝶和真正的捕食者之间有本质的区别,就像 AI 生成的语言和真正的人类交流之间有本质的区别一样。

有人可能会说,使用生成式 AI 来帮助写作是在从 AI 训练的文本中汲取灵感。但这与我们通常所说的一个作家从另一个作家那里汲取灵感是完全不同的。想象一个大学生交上一篇完全由一本书中的长篇引用组成的论文,声称这段引用完美地表达了她想说的话。即使这个学生对此完全坦白,我们也不能说她是在从这本书中汲取灵感。大型语言模型能够重新措辞使源头难以识别,但这并不改变其本质上是在复制而非创造的事实。

正如语言学家 Emily M. Bender 指出的,教师让学生写论文不是因为世界需要更多的学生论文。写论文的目的是培养学生的批判性思维能力。就像举重对各种运动都有益一样,写作能力对学生未来的各种工作都是必要的。使用 ChatGPT 完成作业就像在健身房里使用叉车;你永远不会通过这种方式提高自己的认知能力。

我们必须承认,并非所有的写作都需要富有创意、情感深度或特别出色;有时候,写作仅仅是为了满足某种存在的需求。这类写作可能服务于其他目的,比如吸引广告点击或满足行政程序要求。当人们面临这种写作任务时,我们很难责怪他们利用各种可用工具来提高效率。

然而,我们需要思考一个问题:如果世界上充斥着更多仅仅是为了应付差事而草草完成的文档,这真的是一种进步吗?我们不能天真地认为,只要拒绝使用大型语言模型,对低质量文本的需求就会消失。事实上,情况可能恰恰相反。我担心,随着我们越来越依赖大型语言模型来满足这些写作需求,这些需求本身可能会变得更加庞大和繁琐。

我们似乎正在进入一个荒谬的时代:有人可能会使用 AI 从一份简单的要点清单生成一份冗长的文档,然后将这份文档发送给另一个人,而这个人又会使用 AI 将这份文档重新压缩成一份要点清单。这整个过程看似高效,实则可能是在浪费时间和资源。我们真的能够认真地将这种循环往复的过程称之为进步吗?

这段话提出了一个深刻的问题:在追求效率的同时,我们是否正在牺牲真正有价值的思考和交流?它提醒我们,尽管技术可以简化许多任务,但我们不应忘记写作的本质目的是传达思想和促进理解。过度依赖 AI 生成内容可能会导致一种表面上的效率,但实际上可能降低了信息的质量和意义。我们需要在利用技术提高效率和保持人类思考与创造力之间找到平衡。

尽管未来可能会出现能够完全模仿人类行为的计算机程序,但这种情况在近几年内不太可能实现,这与许多人工智能公司的宣传相悖。即使在那些与创造力无关的领域,当前的人工智能程序仍然存在严重的局限性,这使我们有理由质疑它们是否真的可以被称为 “ 智能 “。

计算机科学家 François Chollet 提出了一个有趣的区分:技能是指完成特定任务的能力,而智能则是快速获取新技能的能力。这个定义很好地反映了我们对人类智能的直观理解。大多数人通过足够的练习都能学会新技能,但学习速度越快的人,我们就认为他越聪明。这个定义的妙处在于它不仅适用于人类,也适用于其他生物;比如当一只狗快速学会新把戏时,我们也会认为这是智能的表现。

2019 年的一项研究中,科学家们教会了老鼠 “ 开车 “。他们将老鼠放在带有三根控制杆的小型塑料容器中,老鼠可以通过触碰这些杆子来控制容器前进或转向。老鼠能看到房间另一端的食物,并试图操控 “ 车辆 “ 朝那个方向移动。经过 24 次每次 5 分钟的训练后,老鼠就掌握了这项在其进化历史上从未遇到过的技能。这无疑是智能的一个很好的例证。

相比之下,现今备受推崇的人工智能程序,如谷歌 DeepMind 开发的 AlphaZero,虽然在国际象棋上超越了人类棋手,但它在训练过程中进行了多达 4400 万局对弈,远远超过任何人一生中可能下的棋局数量。要掌握一个新的游戏,它需要进行同样海量的训练。按照 Chollet 的定义,像 AlphaZero 这样的程序虽然技能很高,但并不特别智能,因为它们获取新技能的效率并不高。目前,如果程序员事先不了解任务的具体信息,还无法编写出能在仅仅 24 次尝试中学会哪怕是一个简单任务的计算机程序。

即使经过数百万英里的训练,自动驾驶汽车仍可能撞上一辆翻倒的拖车,因为这种罕见情况在其训练数据中并不常见。相比之下,即使是第一次上驾驶课的人类学员也会知道在这种情况下应该停车。

尽管生成式人工智能多年来备受关注,但它显著提高经济生产力的能力仍然停留在理论阶段。(高盛最近发布的一份报告就质疑了生成式 AI 的投入产出比。)事实上,生成式 AI 最大的 “ 成就 “ 可能是降低了我们的期望值,无论是对我们阅读的内容,还是对我们自己写作的要求。这种技术从根本上说是去人性化的,因为它将我们简化为不如我们本来面目的存在:意义的创造者和理解者。它减少了世界上的意图和目的性。

有人为大型语言模型辩护,认为人类所说或写的大部分内容本来就不怎么原创。这话虽然没错,但并不重要。当有人对你说 “ 对不起 “ 时,重要的不是这句话有多么独特或罕见,而是说这句话的人的真诚态度。同样,当你告诉某人你很高兴见到他们时,即使这句话听起来很普通,但因为是你在特定情境下说的,所以它就是有意义的。

这个道理同样适用于艺术创作。无论你是在创作小说、绘画还是电影,你都在与观众进行一种独特的交流。你的作品不需要与历史上的每一件艺术品都截然不同才有价值。 恰恰相反,正是因为这是你基于自己独特的生活经历所创作的,并且在观众生命中的特定时刻被接收到,它才显得新颖和有意义。我们都是前人智慧的继承者,但正是通过我们与他人的互动和交流,我们才能为这个世界带来新的意义。这是任何自动完成算法都无法做到的,不要让任何人告诉你相反的话。


Why A.I. Isn’t Going to Make Art

To create a novel or a painting, an artist makes choices that are fundamentally alien to artificial intelligence.

By Ted Chiang

August 31, 2024

In 1953, Roald Dahl published “The Great Automatic Grammatizator,” a short story about an electrical engineer who secretly desires to be a writer. One day, after completing construction of the world’s fastest calculating machine, the engineer realizes that “English grammar is governed by rules that are almost mathematical in their strictness.” He constructs a fiction-writing machine that can produce a five-thousand-word short story in thirty seconds; a novel takes fifteen minutes and requires the operator to manipulate handles and foot pedals, as if he were driving a car or playing an organ, to regulate the levels of humor and pathos. The resulting novels are so popular that, within a year, half the fiction published in English is a product of the engineer’s invention.

Is there anything about art that makes us think it can’t be created by pushing a button, as in Dahl’s imagination? Right now, the fiction generated by large language models like ChatGPT is terrible, but one can imagine that such programs might improve in the future. How good could they get? Could they get better than humans at writing fiction—or making paintings or movies—in the same way that calculators are better at addition and subtraction?

Art is notoriously hard to define, and so are the differences between good art and bad art. But let me offer a generalization: art is something that results from making a lot of choices. This might be easiest to explain if we use fiction writing as an example. When you are writing fiction, you are—consciously or unconsciously—making a choice about almost every word you type; to oversimplify, we can imagine that a ten-thousand-word short story requires something on the order of ten thousand choices. When you give a generative-A.I. program a prompt, you are making very few choices; if you supply a hundred-word prompt, you have made on the order of a hundred choices.

If an A.I. generates a ten-thousand-word story based on your prompt, it has to fill in for all of the choices that you are not making. There are various ways it can do this. One is to take an average of the choices that other writers have made, as represented by text found on the Internet; that average is equivalent to the least interesting choices possible, which is why A.I.-generated text is often really bland. Another is to instruct the program to engage in style mimicry, emulating the choices made by a specific writer, which produces a highly derivative story. In neither case is it creating interesting art.

I think the same underlying principle applies to visual art, although it’s harder to quantify the choices that a painter might make. Real paintings bear the mark of an enormous number of decisions. By comparison, a person using a text-to-image program like DALL-E enters a prompt such as “A knight in a suit of armor fights a fire-breathing dragon,” and lets the program do the rest. (The newest version of DALL-E accepts prompts of up to four thousand characters—hundreds of words, but not enough to describe every detail of a scene.) Most of the choices in the resulting image have to be borrowed from similar paintings found online; the image might be exquisitely rendered, but the person entering the prompt can’t claim credit for that.

Some commentators imagine that image generators will affect visual culture as much as the advent of photography once did. Although this might seem superficially plausible, the idea that photography is similar to generative A.I. deserves closer examination. When photography was first developed, I suspect it didn’t seem like an artistic medium because it wasn’t apparent that there were a lot of choices to be made; you just set up the camera and start the exposure. But over time people realized that there were a vast number of things you could do with cameras, and the artistry lies in the many choices that a photographer makes. It might not always be easy to articulate what the choices are, but when you compare an amateur’s photos to a professional’s, you can see the difference. So then the question becomes: Is there a similar opportunity to make a vast number of choices using a text-to-image generator? I think the answer is no. An artist—whether working digitally or with paint—implicitly makes far more decisions during the process of making a painting than would fit into a text prompt of a few hundred words.

We can imagine a text-to-image generator that, over the course of many sessions, lets you enter tens of thousands of words into its text box to enable extremely fine-grained control over the image you’re producing; this would be something analogous to Photoshop with a purely textual interface. I’d say that a person could use such a program and still deserve to be called an artist. The film director Bennett Miller has used DALL-E 2 to generate some very striking images that have been exhibited at the Gagosian gallery; to create them, he crafted detailed text prompts and then instructed DALL-E to revise and manipulate the generated images again and again. He generated more than a hundred thousand images to arrive at the twenty images in the exhibit. But he has said that he hasn’t been able to obtain comparable results on later releases of DALL-E. I suspect this might be because Miller was using DALL-E for something it’s not intended to do; it’s as if he hacked Microsoft Paint to make it behave like Photoshop, but as soon as a new version of Paint was released, his hacks stopped working. OpenAI probably isn’t trying to build a product to serve users like Miller, because a product that requires a user to work for months to create an image isn’t appealing to a wide audience. The company wants to offer a product that generates images with little effort.

It’s harder to imagine a program that, over many sessions, helps you write a good novel. This hypothetical writing program might require you to enter a hundred thousand words of prompts in order for it to generate an entirely different hundred thousand words that make up the novel you’re envisioning. It’s not clear to me what such a program would look like. Theoretically, if such a program existed, the user could perhaps deserve to be called the author. But, again, I don’t think companies like OpenAI want to create versions of ChatGPT that require just as much effort from users as writing a novel from scratch. The selling point of generative A.I. is that these programs generate vastly more than you put into them, and that is precisely what prevents them from being effective tools for artists.

The companies promoting generative-A.I. programs claim that they will unleash creativity. In essence, they are saying that art can be all inspiration and no perspiration—but these things cannot be easily separated. I’m not saying that art has to involve tedium. What I’m saying is that art requires making choices at every scale; the countless small-scale choices made during implementation are just as important to the final product as the few large-scale choices made during the conception. It is a mistake to equate “large-scale” with “important” when it comes to the choices made when creating art; the interrelationship between the large scale and the small scale is where the artistry lies.

Believing that inspiration outweighs everything else is, I suspect, a sign that someone is unfamiliar with the medium. I contend that this is true even if one’s goal is to create entertainment rather than high art. People often underestimate the effort required to entertain; a thriller novel may not live up to Kafka’s ideal of a book—an “axe for the frozen sea within us”—but it can still be as finely crafted as a Swiss watch. And an effective thriller is more than its premise or its plot. I doubt you could replace every sentence in a thriller with one that is semantically equivalent and have the resulting novel be as entertaining. This means that its sentences—and the small-scale choices they represent—help to determine the thriller’s effectiveness.

Many novelists have had the experience of being approached by someone convinced that they have a great idea for a novel, which they are willing to share in exchange for a fifty-fifty split of the proceeds. Such a person inadvertently reveals that they think formulating sentences is a nuisance rather than a fundamental part of storytelling in prose. Generative A.I. appeals to people who think they can express themselves in a medium without actually working in that medium. But the creators of traditional novels, paintings, and films are drawn to those art forms because they see the unique expressive potential that each medium affords. It is their eagerness to take full advantage of those potentialities that makes their work satisfying, whether as entertainment or as art.

Of course, most pieces of writing, whether articles or reports or e-mails, do not come with the expectation that they embody thousands of choices. In such cases, is there any harm in automating the task? Let me offer another generalization: any writing that deserves your attention as a reader is the result of effort expended by the person who wrote it. Effort during the writing process doesn’t guarantee the end product is worth reading, but worthwhile work cannot be made without it. The type of attention you pay when reading a personal e-mail is different from the type you pay when reading a business report, but in both cases it is only warranted when the writer put some thought into it.

Recently, Google aired a commercial during the Paris Olympics for Gemini, its competitor to OpenAI’s GPT-4. The ad shows a father using Gemini to compose a fan letter, which his daughter will send to an Olympic athlete who inspires her. Google pulled the commercial after widespread backlash from viewers; a media professor called it “one of the most disturbing commercials I’ve ever seen.” It’s notable that people reacted this way, even though artistic creativity wasn’t the attribute being supplanted. No one expects a child’s fan letter to an athlete to be extraordinary; if the young girl had written the letter herself, it would likely have been indistinguishable from countless others. The significance of a child’s fan letter—both to the child who writes it and to the athlete who receives it—comes from its being heartfelt rather than from its being eloquent.

Many of us have sent store-bought greeting cards, knowing that it will be clear to the recipient that we didn’t compose the words ourselves. We don’t copy the words from a Hallmark card in our own handwriting, because that would feel dishonest. The programmer Simon Willison has described the training for large language models as “money laundering for copyrighted data,” which I find a useful way to think about the appeal of generative-A.I. programs: they let you engage in something like plagiarism, but there’s no guilt associated with it because it’s not clear even to you that you’re copying.

Some have claimed that large language models are not laundering the texts they’re trained on but, rather, learning from them, in the same way that human writers learn from the books they’ve read. But a large language model is not a writer; it’s not even a user of language. Language is, by definition, a system of communication, and it requires an intention to communicate. Your phone’s auto-complete may offer good suggestions or bad ones, but in neither case is it trying to say anything to you or the person you’re texting. The fact that ChatGPT can generate coherent sentences invites us to imagine that it understands language in a way that your phone’s auto-complete does not, but it has no more intention to communicate.

It is very easy to get ChatGPT to emit a series of words such as “I am happy to see you.” There are many things we don’t understand about how large language models work, but one thing we can be sure of is that ChatGPT is not happy to see you. A dog can communicate that it is happy to see you, and so can a prelinguistic child, even though both lack the capability to use words. ChatGPT feels nothing and desires nothing, and this lack of intention is why ChatGPT is not actually using language. What makes the words “I’m happy to see you” a linguistic utterance is not that the sequence of text tokens that it is made up of are well formed; what makes it a linguistic utterance is the intention to communicate something.

Because language comes so easily to us, it’s easy to forget that it lies on top of these other experiences of subjective feeling and of wanting to communicate that feeling. We’re tempted to project those experiences onto a large language model when it emits coherent sentences, but to do so is to fall prey to mimicry; it’s the same phenomenon as when butterflies evolve large dark spots on their wings that can fool birds into thinking they’re predators with big eyes. There is a context in which the dark spots are sufficient; birds are less likely to eat a butterfly that has them, and the butterfly doesn’t really care why it’s not being eaten, as long as it gets to live. But there is a big difference between a butterfly and a predator that poses a threat to a bird.

A person using generative A.I. to help them write might claim that they are drawing inspiration from the texts the model was trained on, but I would again argue that this differs from what we usually mean when we say one writer draws inspiration from another. Consider a college student who turns in a paper that consists solely of a five-page quotation from a book, stating that this quotation conveys exactly what she wanted to say, better than she could say it herself. Even if the student is completely candid with the instructor about what she’s done, it’s not accurate to say that she is drawing inspiration from the book she’s citing. The fact that a large language model can reword the quotation enough that the source is unidentifiable doesn’t change the fundamental nature of what’s going on.

As the linguist Emily M. Bender has noted, teachers don’t ask students to write essays because the world needs more student essays. The point of writing essays is to strengthen students’ critical-thinking skills; in the same way that lifting weights is useful no matter what sport an athlete plays, writing essays develops skills necessary for whatever job a college student will eventually get. Using ChatGPT to complete assignments is like bringing a forklift into the weight room; you will never improve your cognitive fitness that way.

Not all writing needs to be creative, or heartfelt, or even particularly good; sometimes it simply needs to exist. Such writing might support other goals, such as attracting views for advertising or satisfying bureaucratic requirements. When people are required to produce such text, we can hardly blame them for using whatever tools are available to accelerate the process. But is the world better off with more documents that have had minimal effort expended on them? It would be unrealistic to claim that if we refuse to use large language models, then the requirements to create low-quality text will disappear. However, I think it is inevitable that the more we use large language models to fulfill those requirements, the greater those requirements will eventually become. We are entering an era where someone might use a large language model to generate a document out of a bulleted list, and send it to a person who will use a large language model to condense that document into a bulleted list. Can anyone seriously argue that this is an improvement?

It’s not impossible that one day we will have computer programs that can do anything a human being can do, but, contrary to the claims of the companies promoting A.I., that is not something we’ll see in the next few years. Even in domains that have absolutely nothing to do with creativity, current A.I. programs have profound limitations that give us legitimate reasons to question whether they deserve to be called intelligent at all.

The computer scientist François Chollet has proposed the following distinction: skill is how well you perform at a task, while intelligence is how efficiently you gain new skills. I think this reflects our intuitions about human beings pretty well. Most people can learn a new skill given sufficient practice, but the faster the person picks up the skill, the more intelligent we think the person is. What’s interesting about this definition is that—unlike I.Q. tests—it’s also applicable to nonhuman entities; when a dog learns a new trick quickly, we consider that a sign of intelligence.

In 2019, researchers conducted an experiment in which they taught rats how to drive. They put the rats in little plastic containers with three copper-wire bars; when the mice put their paws on one of these bars, the container would either go forward, or turn left or turn right. The rats could see a plate of food on the other side of the room and tried to get their vehicles to go toward it. The researchers trained the rats for five minutes at a time, and after twenty-four practice sessions, the rats had become proficient at driving. Twenty-four trials were enough to master a task that no rat had likely ever encountered before in the evolutionary history of the species. I think that’s a good demonstration of intelligence.

Now consider the current A.I. programs that are widely acclaimed for their performance. AlphaZero, a program developed by Google’s DeepMind, plays chess better than any human player, but during its training it played forty-four million games, far more than any human can play in a lifetime. For it to master a new game, it will have to undergo a similarly enormous amount of training. By Chollet’s definition, programs like AlphaZero are highly skilled, but they aren’t particularly intelligent, because they aren’t efficient at gaining new skills. It is currently impossible to write a computer program capable of learning even a simple task in only twenty-four trials, if the programmer is not given information about the task beforehand.

Self-driving cars trained on millions of miles of driving can still crash into an overturned trailer truck, because such things are not commonly found in their training data, whereas humans taking their first driving class will know to stop. More than our ability to solve algebraic equations, our ability to cope with unfamiliar situations is a fundamental part of why we consider humans intelligent. Computers will not be able to replace humans until they acquire that type of competence, and that is still a long way off; for the time being, we’re just looking for jobs that can be done with turbocharged auto-complete.

Despite years of hype, the ability of generative A.I. to dramatically increase economic productivity remains theoretical. (Earlier this year, Goldman Sachs released a report titled “Gen AI: Too Much Spend, Too Little Benefit?”) The task that generative A.I. has been most successful at is lowering our expectations, both of the things we read and of ourselves when we write anything for others to read. It is a fundamentally dehumanizing technology because it treats us as less than what we are: creators and apprehenders of meaning. It reduces the amount of intention in the world.

Some individuals have defended large language models by saying that most of what human beings say or write isn’t particularly original. That is true, but it’s also irrelevant. When someone says “I’m sorry” to you, it doesn’t matter that other people have said sorry in the past; it doesn’t matter that “I’m sorry” is a string of text that is statistically unremarkable. If someone is being sincere, their apology is valuable and meaningful, even though apologies have previously been uttered. Likewise, when you tell someone that you’re happy to see them, you are saying something meaningful, even if it lacks novelty.

Something similar holds true for art. Whether you are creating a novel or a painting or a film, you are engaged in an act of communication between you and your audience. What you create doesn’t have to be utterly unlike every prior piece of art in human history to be valuable; the fact that you’re the one who is saying it, the fact that it derives from your unique life experience and arrives at a particular moment in the life of whoever is seeing your work, is what makes it new. We are all products of what has come before us, but it’s by living our lives in interaction with others that we bring meaning into the world. That is something that an auto-complete algorithm can never do, and don’t let anyone tell you otherwise.

本文由作者按照 CC BY 4.0 进行授权