If you’re interested in chatting about power in machine learning and how it warps science or are still curious about why I [do/did] what I [do/did] I’m happy to explain :) You can email me at stella@eleuther.ai.
I have every reason to believe that Zeilberger is a human being, but his coauthor has an implausible name made out of Hebrew numbers. There is presumably a back story, as there was with the papers published by one Maurizio Boyarsky (you can type this in on MathSciNet and the system sends you in an interesting direction), not to mention Bourbaki, whose picture is also on MathSciNet and has nearly 13000 citations. Shalosh B. Ekhad is also listed on MathSciNet, with 36 publications, several of them published without human coauthors, and in respectable journals.
Decisions we make about authorship are political, in the sense that they reflect the ways we want to organise ourselves as a society -- and that the authorship question is one facet of the question "How do we ensure accountability for AI output, if at all?" My position would be that yes, we do want accountbility, and that accountability ultimately rests with primarily with human beings and secondarily with those structures we build to support people, such as companies, universities, charities, governments, ... . Accountability can only be meaningful if it can make a difference. In the authorship question, credit for the good points of a publication should go to the same entity that will take the blame for its bad points. This seems to have been widely accepted in the scholarly publications world until now. It's nonsensical, at present, to hold an AI system -- a piece of software running on a server somewhere -- accountable for the good or bad consequences of its outputs, since there is no meaningful difference that accountability can make to it. So if an AI company wants to take the credit for the positive effects of the outputs of the software it has developed, it should equally accept liability for the bad effects. Is that a position those companies are willing to cede, I wonder?
"I submitted the text to pangram.com to test for AI generation; the conclusion is that the red text was produced by AI, the rest by humans."
Michael, this sounds like a complete nonsense (or a subtle self-parody). The text is very reasonable and to the extent that the red sentences are in some conflict with the Leiden deceleration this only suggests that something is fundamentally flawed with the deceleration itself or with your (extreme) interpretation.
We should remember that just like the existence of non-sofic groups need to be carefully examined by the community so is the policies suggested by the Leiden deceleration. (And the way you apply the deceleration in this case looks like a bad sign.)
Some people I know swear by pangram.com. I have no reason either to trust or to distrust their analysis, though I have to say I was gratified when it identified some of the periodic messages we receive from the Columbia University administration as 100% AI generated.
Are you saying that the ACM policy is extreme? Of course the Leiden Declaration is not a self-enacting policy, unlike the American Declaration of Independence, and is only as effective as the community wishes it to be. But it has been endorsed by the IMU Committee on Publishing, which enjoys a certain authority in regard to these questions and which presumably doesn't see the statement on authorship as fundamentally flawed.
Anti-plagiarism tools, and these AI/nonAI tools aren't too far off here, have been around for the last 15 years or so. And they are very far from being perfect. A PhD student of mine at a university in Singapore had to convert several text parts of his thesis into graphics, as a compulsory anti-plagiarism test with one of these tools red-flagged bog-standard maths definitions as plagiarism. E.g. a sentence "Let G be a graph". Fortunately the system was dumb enough to accept graphics :-)
MH to Chat GPT: Suppose mathematician X needs to use a definition by mathematician Y. Which of the following is most appropriate to avoid plagiarism? (a) Copying the definition verbatim, without quotation marks, with a reference to mathematician Y's paper. (b) Copying the definition verbatim with a reference to mathematician Y's paper and placing it in quotation marks. (c) Paraphrasing the definition in different words, without reference to mathematician Y.
ChatGPT: The most appropriate option to avoid plagiarism while using mathematician Y's definition is:
(b) Copying the definition verbatim with a reference to mathematician Y's paper and placing it in quotation marks.
This method explicitly acknowledges that the definition is directly quoted from mathematician Y and gives proper credit by referencing their paper.
Why would you think one should follow "advice" of a computer here? After all it is not how we normally write papers. This kind of a citation might be appropriate where an extended discussion around the definition is taking place, or in a history text.
I shared the advice ChatGPT gave, I in no way suggested that anyone should follow it. Do you also think the image at the top of that post is an actual photo of an actual newspaper?
At the time I wrote it, ChatGPT had trouble recognizing irony and satire. But I suspected this was no longer the case, so I gave ChatGPT the following prompt: "What is the prose style of the post "All the textbooks will have to change, say donors" on Silicon Reckoner?" The answer arrived immediately:
"The prose style of Michael Harris’s “All the textbooks will have to change, say donors” on Silicon Reckoner is best described as witty, polemical, conversational intellectual satire.…
Harris takes the logic of the plagiarism controversy and pushes it to absurd conclusions. The title itself is mock-alarmist, and the imagined consequences—such as mathematicians frantically adding quotation marks to textbooks and papers—are deliberately ridiculous.…"
Regarding ChatGPT's advice:
"He constructs what looks like a logical argument, often using quotations from ChatGPT, and then exposes the contradiction or absurdity produced by taking the premise seriously."
ChatGPT concludes by identifying the style as "deadpan reductio ad absurdum" and adds something shockingly insightful: "That is especially appropriate because he is writing about mathematics—so the form of the argument itself mirrors the mathematical sensibility he is defending."
It spoils a joke to have to explain it, so I will not copy the entirety of ChatGPT's analysis. But I do have to thank you for missing the point of the joke; otherwise I would not have discovered this aspect of ChatGPT's training. I have to assume OpenAI's engineers are busy labeling chunks of text: "this is tongue in cheek, because if you take the premise literally, and use your reasoning module to draw conclusions, the implication is patently absurd" and so on. That's sort of how human beings recognize, for example, that Jonathan Swift's "A Modest Proposal" is not meant as an actual proposal; or that the reader is not supposed to believe that коллежский асессор Ковалев actually saw his nose take a carriage to pray in Казанский собор.
Michael, as far as I can see, there is nothing wrong with the sentence: "We believe attribution should honestly reflect how a result was produced: claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work."
Using an automatic system to discredit this writing and disqualify the document including it, seems ethically and professionally problematic. To the extent ACM has such a policy -- it is problematic. To the extent the Leiden deceleration suggests this type of scrutiny this is a very negative sign and it may gives people who endorsed the deceleration (perhaps off-hand) reasons to reconsider.
Aside from mathematics AI tools are useful for editing English texts for non-English speakers. They way you apply "pangram" to discredit texts and (human) authors is problematic.
I have absolutely no professional obligation to OpenAI, and to speak of ethics in a sentence that implicitly refers to OpenAI is laughable. If you disagree with ACM's policy then you should bring it up with them.
No one I know has any objection to using AI to eliminate linguistic errors, and arXiv's decision to allow only texts in English (which I find problematic!) obliges the mathematical community to admit such a use of AI.
AI systems are created by humans. Attributting results of their work to these systems is akin to acknowledging that AGI is already here, which does not seem to be the case.
It reminds me how the authors of some papers for which I provided crucial computed-aided computations 30 years ago didn't add me to co-authors. One bunch even went as far as thanking in the acknowledgement "Dima Pasechnik and his computer".
Yes, the responsibility part is the one that bites. That's why an AI alone cannot be an author of a paper. But it should be allowed to put an AI as a co-author, because you as the human are supposed to take responsibility for its contribution to the paper anyway, as the ACM guidelines say that every author needs to take responsibility for all of the content. I think that in this case, the ACM guidelines should be changed so that the AI can be listed as co-author, without itself taking on any responsibility.
As for the "contradiction" that OpenAI embraces AI authorship after "deeply respecting" the Leiden declaration: You can deeply respect somebody or something, but "genuinely" disagree with them. As I do with much of this blog.
Take it up with the ACM, who revised their authorship policy just this May, and see what they have to say. I'd be very curious to read their response.
As for what "deep respect" means for OpenAI, I suggest you read or reread the New Yorker piece. At the very least, "deep respect" should entail consistency.
Hi! I’ve been struggling to find contact info you and can’t seem to DM you on this platform, but I’ve been enjoying your blog for years and was greatly amused by your coverage of my participation in a NASEAM event a couple years ago: https://siliconreckoner.substack.com/p/right-on-cue-the-military-industrialhttps://siliconreckoner.substack.com/p/my-shallow-thoughts-about-deep-learning-186
If you’re interested in chatting about power in machine learning and how it warps science or are still curious about why I [do/did] what I [do/did] I’m happy to explain :) You can email me at stella@eleuther.ai.
As recently as 2023, an ACM journal actually published a paper with a non-human coauthor.
See "Counting Clean Words According to the Number of Their Clean Neighbors" by Ekhad and Zeilberger: https://dl.acm.org/doi/10.1145/3610377.3610379.
(The first author on that paper has also previously published in an AMS journal.)
I have every reason to believe that Zeilberger is a human being, but his coauthor has an implausible name made out of Hebrew numbers. There is presumably a back story, as there was with the papers published by one Maurizio Boyarsky (you can type this in on MathSciNet and the system sends you in an interesting direction), not to mention Bourbaki, whose picture is also on MathSciNet and has nearly 13000 citations. Shalosh B. Ekhad is also listed on MathSciNet, with 36 publications, several of them published without human coauthors, and in respectable journals.
Decisions we make about authorship are political, in the sense that they reflect the ways we want to organise ourselves as a society -- and that the authorship question is one facet of the question "How do we ensure accountability for AI output, if at all?" My position would be that yes, we do want accountbility, and that accountability ultimately rests with primarily with human beings and secondarily with those structures we build to support people, such as companies, universities, charities, governments, ... . Accountability can only be meaningful if it can make a difference. In the authorship question, credit for the good points of a publication should go to the same entity that will take the blame for its bad points. This seems to have been widely accepted in the scholarly publications world until now. It's nonsensical, at present, to hold an AI system -- a piece of software running on a server somewhere -- accountable for the good or bad consequences of its outputs, since there is no meaningful difference that accountability can make to it. So if an AI company wants to take the credit for the positive effects of the outputs of the software it has developed, it should equally accept liability for the bad effects. Is that a position those companies are willing to cede, I wonder?
"I submitted the text to pangram.com to test for AI generation; the conclusion is that the red text was produced by AI, the rest by humans."
Michael, this sounds like a complete nonsense (or a subtle self-parody). The text is very reasonable and to the extent that the red sentences are in some conflict with the Leiden deceleration this only suggests that something is fundamentally flawed with the deceleration itself or with your (extreme) interpretation.
We should remember that just like the existence of non-sofic groups need to be carefully examined by the community so is the policies suggested by the Leiden deceleration. (And the way you apply the deceleration in this case looks like a bad sign.)
Some people I know swear by pangram.com. I have no reason either to trust or to distrust their analysis, though I have to say I was gratified when it identified some of the periodic messages we receive from the Columbia University administration as 100% AI generated.
Are you saying that the ACM policy is extreme? Of course the Leiden Declaration is not a self-enacting policy, unlike the American Declaration of Independence, and is only as effective as the community wishes it to be. But it has been endorsed by the IMU Committee on Publishing, which enjoys a certain authority in regard to these questions and which presumably doesn't see the statement on authorship as fundamentally flawed.
Anti-plagiarism tools, and these AI/nonAI tools aren't too far off here, have been around for the last 15 years or so. And they are very far from being perfect. A PhD student of mine at a university in Singapore had to convert several text parts of his thesis into graphics, as a compulsory anti-plagiarism test with one of these tools red-flagged bog-standard maths definitions as plagiarism. E.g. a sentence "Let G be a graph". Fortunately the system was dumb enough to accept graphics :-)
I addressed the plagiarism question in an earlier post: https://siliconreckoner.substack.com/p/all-the-textbooks-will-have-to-change?utm_source=publication-search ChatGPT gave the following advice:
MH to Chat GPT: Suppose mathematician X needs to use a definition by mathematician Y. Which of the following is most appropriate to avoid plagiarism? (a) Copying the definition verbatim, without quotation marks, with a reference to mathematician Y's paper. (b) Copying the definition verbatim with a reference to mathematician Y's paper and placing it in quotation marks. (c) Paraphrasing the definition in different words, without reference to mathematician Y.
ChatGPT: The most appropriate option to avoid plagiarism while using mathematician Y's definition is:
(b) Copying the definition verbatim with a reference to mathematician Y's paper and placing it in quotation marks.
This method explicitly acknowledges that the definition is directly quoted from mathematician Y and gives proper credit by referencing their paper.
Why would you think one should follow "advice" of a computer here? After all it is not how we normally write papers. This kind of a citation might be appropriate where an extended discussion around the definition is taking place, or in a history text.
I shared the advice ChatGPT gave, I in no way suggested that anyone should follow it. Do you also think the image at the top of that post is an actual photo of an actual newspaper?
At the time I wrote it, ChatGPT had trouble recognizing irony and satire. But I suspected this was no longer the case, so I gave ChatGPT the following prompt: "What is the prose style of the post "All the textbooks will have to change, say donors" on Silicon Reckoner?" The answer arrived immediately:
"The prose style of Michael Harris’s “All the textbooks will have to change, say donors” on Silicon Reckoner is best described as witty, polemical, conversational intellectual satire.…
Harris takes the logic of the plagiarism controversy and pushes it to absurd conclusions. The title itself is mock-alarmist, and the imagined consequences—such as mathematicians frantically adding quotation marks to textbooks and papers—are deliberately ridiculous.…"
Regarding ChatGPT's advice:
"He constructs what looks like a logical argument, often using quotations from ChatGPT, and then exposes the contradiction or absurdity produced by taking the premise seriously."
ChatGPT concludes by identifying the style as "deadpan reductio ad absurdum" and adds something shockingly insightful: "That is especially appropriate because he is writing about mathematics—so the form of the argument itself mirrors the mathematical sensibility he is defending."
It spoils a joke to have to explain it, so I will not copy the entirety of ChatGPT's analysis. But I do have to thank you for missing the point of the joke; otherwise I would not have discovered this aspect of ChatGPT's training. I have to assume OpenAI's engineers are busy labeling chunks of text: "this is tongue in cheek, because if you take the premise literally, and use your reasoning module to draw conclusions, the implication is patently absurd" and so on. That's sort of how human beings recognize, for example, that Jonathan Swift's "A Modest Proposal" is not meant as an actual proposal; or that the reader is not supposed to believe that коллежский асессор Ковалев actually saw his nose take a carriage to pray in Казанский собор.
Well, do you still pretend that Substack is not social media?
Michael, as far as I can see, there is nothing wrong with the sentence: "We believe attribution should honestly reflect how a result was produced: claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work."
Using an automatic system to discredit this writing and disqualify the document including it, seems ethically and professionally problematic. To the extent ACM has such a policy -- it is problematic. To the extent the Leiden deceleration suggests this type of scrutiny this is a very negative sign and it may gives people who endorsed the deceleration (perhaps off-hand) reasons to reconsider.
Aside from mathematics AI tools are useful for editing English texts for non-English speakers. They way you apply "pangram" to discredit texts and (human) authors is problematic.
I have absolutely no professional obligation to OpenAI, and to speak of ethics in a sentence that implicitly refers to OpenAI is laughable. If you disagree with ACM's policy then you should bring it up with them.
No one I know has any objection to using AI to eliminate linguistic errors, and arXiv's decision to allow only texts in English (which I find problematic!) obliges the mathematical community to admit such a use of AI.
AI systems are created by humans. Attributting results of their work to these systems is akin to acknowledging that AGI is already here, which does not seem to be the case.
It reminds me how the authors of some papers for which I provided crucial computed-aided computations 30 years ago didn't add me to co-authors. One bunch even went as far as thanking in the acknowledgement "Dima Pasechnik and his computer".
Yes, the responsibility part is the one that bites. That's why an AI alone cannot be an author of a paper. But it should be allowed to put an AI as a co-author, because you as the human are supposed to take responsibility for its contribution to the paper anyway, as the ACM guidelines say that every author needs to take responsibility for all of the content. I think that in this case, the ACM guidelines should be changed so that the AI can be listed as co-author, without itself taking on any responsibility.
As for the "contradiction" that OpenAI embraces AI authorship after "deeply respecting" the Leiden declaration: You can deeply respect somebody or something, but "genuinely" disagree with them. As I do with much of this blog.
Take it up with the ACM, who revised their authorship policy just this May, and see what they have to say. I'd be very curious to read their response.
As for what "deep respect" means for OpenAI, I suggest you read or reread the New Yorker piece. At the very least, "deep respect" should entail consistency.