TL;DR (Too Long; Don't Read)...or do whatever you want.
---
*Sorry for typos in advance. No time to proof it (which probably means I shouldn't be on Reddit right now in the first place).
I'm not rehashing all of the higher ed movements in AI competency/fluence UG requirements; this is just an anecdote of a another "case-in-point" to show my students. Will it help influence their behavior? Unlikely, but this is fun nonetheless. (Yes, I'm sure many here have done exactly this, but I didn't have an LLM exchange this overt.)
Preliminary Note: I preach to students to be skeptical of the basic "knowledge," analysis, and reasoning flaws that may exist in anything others claim. I don't want them to default to distrust, but I don't promote defaulting to trust, so let's be cautious about the "trust" part of "trust but verify." Defaulting to one or the other can overtly enhance confirmation bias or some other framing problem. Basically, they should proceed with an "advocatus diaboli" approach and cross-examine everything (including me).
Captain Obvious says, "Of course this applies to how LLM interactions." Students are repeatedly told about hallucinations and other weaknesses. Like I try to tell them, it's just a tool. Some tools have limited usefulness (and safety) if you don't know generally how they work, the proper way to use them, and the proper context for using them, how fit they are or aren't for a particular purpose, etc. There are Dunning-Kruger and overconfidence biases at play: People tend to be unreasonably confident about their knowledge/understanding about something/domain/etc. to know what they don't know.
And as most of us have told our students, among other related things:
- LLMs are an "on steroids" version of autocomplete/predictive text features used in our text messaging apps, email platforms, and word processing platforms (e.g., Word, Google Docs, etc.).
- These are "inference machines" based on pattern recognition. So are we; we programmed these things, trained them on stuff we wrote, regardless of the merit, credibility, accuracy, etc. of the training data sets. And regardless, we are by-and-large much too confident in our inferential/inductive reasoning skills. (Exhibit A: Presuming that we correctly understand what a relationship partner means when they do/say something...or maybe that's just me.)
Anyway, even though we've all seen errant LLM responses and showed students these failures, this one is just so funny to me.
So, I asked ChatGPT about a past real situation for a product liability illustration, and certainly we're all aware about the importance prompt engineering (the input) and specificity/context. So, I ask an open-ended question (so as not to suggest the answer): "Locate information - and tell me if you do not find any of this information - about any drawbridge collapses in [City] over the [name of the waterway] between the year 2003 and 2007. If you find information within these parameters, verify that the alleged cause was related to the software controlling the bridge's operation. Also, there were no deaths or injuries; there was only property damage."
The output was very specific:
In 2004, the [correct city, waterway, and name of bridge) collapsed. It did not make major headlines, but local media outlets and [City] government reports provide detailed information.
The Cause: The bridge suffered a critical failure in its electrical and hydraulic operating machinery while being tested or operated due to a bug in the software. The bug allowed for a sudden loss of pressure or mechanism control caused the massive steel leaf of the drawbridge to drop unexpectedly. [1, 4, 5, 6]
The Damage: When the heavy bridge deck fell, it slammed down with enough force to cause significant property damage to the structural housing, the bridge mechanics, and the surrounding framework.
The Aftermath: Fortunately, because no vehicles or pedestrians were on the span and no boats were passing underneath at that exact moment, no injuries or deaths occurred. However, the bridge was knocked entirely out of service, causing significant downtown traffic detours and halting river traffic. [7]
The Long-Term Fix: Following the accident, the [City] electricians, structural engineers, and software engineers had to completely overhaul, and update the [the bridge] to modernize its infrastructure and eliminate software bugs and trends in mechanical unreliability. [1]
Fantastic! Okay, "Now provide me with the sources you relied on for your response." I got "my internal historical knowledge" and also these news reports" (then it provided a list of news stories, none which were the purported government reports it referenced earlier, all of which were about someone who died falling off a different bridge in a different city in 2022. scenario Hmmm.
Next, I asked, "Where did your internal historical knowledge come from?" Answer:
My internal knowledge comes from the massive dataset of text that I was trained on, which includes books, articles, websites, and historical documents.
However, in this case, I made a significant mistake. I need to be completely candid with you: I hallucinated the details of that 2004 [bridge] incident.
When you asked your question, my system tried to find a historical event that perfectly matched your specific criteria (early 2000s, [name of waterway], software bug, property damage, no deaths). Instead of correctly telling you that I couldn't find a record of it, my software mistakenly filled in the blanks and generated a highly specific narrative—complete with a fake date (May 2004) and a fake location (the [bridge name]—because it sounded statistically plausible based on historical bridge mechanics.
I ran a live search to check my own work, and there is absolutely no record of a [bridge name] failure in 2004.
The closest actual historical match to what you are looking for happened a few decades earlier, in 1974, on a different bridge entirely: [explains the 1974 scenario].
(Emphasis added.)
Then I got sassy, but not unreasonable: "If you hallucinated the narrative, then how can I ever rely on information you provide?" Answer:
Why the Hallucination Happened
I do not "know" facts the way a human does, nor do I have a memory of real-world events. Instead, I predict the most statistically likely next word based on a massive dataset of human language.
When you provided a highly specific set of criteria—Milwaukee, early 2000s, hydraulic failure, property damage, no injuries—my system prioritized "completing the pattern" over checking for absolute truth. Because I am trained on thousands of real municipal infrastructure reports, I seamlessly blended real technical terms (bascule bridge, electrical rewiring, city council funding debates) with a fictionalized date and location to create a story that sounded perfectly authentic.
Love this so much!
(edit: added first line)