Khoury News
“Don’t outsource ethics”: Researchers urge caution as AI enters newsrooms
In the fast-paced, information-dense world of professional journalism, AI is quickly becoming a valuable tool. But a closer look shows that chatbots aren't always as neutral as they purport to be.
Khoury PhD candidate Asteria Kaeberlein noticed a shift happening in the world of journalism: the growing use of large language models, or LLMs. The New York Times, for example, launched an AI Initiatives Team in 2024.
So, working with her advisor, Assistant Professor Malihe Alikhani, Kaeberlein researched the use of generative AI tools in journalistic writing and their effects on perceived bias, all as the technology changed in real time. After comparing human-authored, AI-generated, and AI-edited articles, the duo found that while AI generally leaned toward neutrality, its neutral tone can produce harms, including the omission of relevant voices and a failure to curb misinformation.
“People were using LLMs in journalism,” Kaeberlein said. “Obviously LLMs are in a lot of other places as well, but this was kind of a big, ‘Hey we should make sure this is reliable and not biased.’”
Using an LLM-based media bias detector developed by the University of Pennsylvania, Kaeberlein and Alikhani evaluated both political and nonpolitical articles based on five factors: racial bias, gender bias, political bias, tone, and factuality.
“Once we had collected annotations, we compared the LLM ratings to ratings from our human participants on the human-written, LLM-edited, and LLM-generated articles,” Kaeberlein said. “We found that humans would typically rate things very differently than the LLM.”
Indeed, the racial bias category, the one in which LLM and human ratings were most aligned, registered only 40% agreement. Article tone, the least aligned, came in at just 20%.
“We found generally worse results across the board, but it varied depending on which different type,” Kaeberlein said. “So the media bias detector is not reliable if you want to know how humans perceive the media rather than LLMs. This is similar to a lot of other AI methods that don’t properly capture how humans interpret things.”
Participant-wise, the researchers included college students in multiple states. Not only did participants’ responses differ from the LLMs’; the process also highlighted a different contrast, where humans took much longer to assess articles.
Kaeberlein and Alikhani also found that LLMs tend to make text more neutral in editing, often watering down articles that took strong positions. Emotionally charged language was replaced by broad summaries, and explicit viewpoints were replaced with detached terms. This behavior can have dangerous consequences, like preventing marginalized voices from being heard.
“While LLMs tend to make text more neutral — how it’s measured in tone — it doesn’t actually remove biases,” Alikhani said. “Instead it can make biases harder to see or more harmful.”
This observation was especially notable in coverage related to human rights. One striking example involved comments by Kristopher Wells, who described a fellow Canadian politician’s “obsession” with the trans community as “beyond weird.” When rewritten by an LLM, the statement became a broader expression of concern about the government’s focus on transgender issues amid challenges in Alberta’s health care system.
“We had two articles that took opposite positions on bigotry in general — sexism, racism, homophobia, and transphobia — and the LLM called all of them arbitrary,” Kaeberlein said. “In contrast to that, LLMs can also take a paper that’s extremely racist or transphobic and remove some of that.”
LLMs also had trouble with nuance, Kaeberlein and Alikhani found. Cultural nuances, linguistic subtleties, and what Alikhani referred to as “mini cultures” — social contexts with distinct cultural practices, references, and ways of communicating — are easy for the technology to miss, both when it’s generating an article and editing human-written text.
“Journalists are not collectors; they report, they produce, and AI is not a producer in that sense,” Alikhani said. “Unless information actually exists in the AI training data for a considerable amount of time and with a good size for training, then AI wouldn’t be able to cover it or help with it.”
Hallucinations, stereotype reinforcement, and sycophantic behavior all remain concerns for LLMs, and can influence perceptions of dependability, according to Kaeberlein.
Research has shown that humans prefer models that exhibit convincing, flattering responses. AI companies, which are incentivized to create successful chatbots, have noticed; when OpenAI released GPT-5 in August 2025, they noted that earlier versions had become so jarringly sycophantic, GPT-5 was engineered to be “less effusively agreeable” and “more subtle and thoughtful in follow-ups.”
“Because LLMs take a more positive tone, and people are more willing to believe them, they can make what is biased seem more trustworthy in a way, even when it’s not,” Kaeberlein said.
Beyond tone, the researchers found that LLMs may also reshape information by introducing interpretations that aren’t in the original text. For instance, when a human-written article reported that Jordan Bardella, a member of the European Parliament, rejected accusations of fostering a climate of hatred against his party, the LLM-generated response added that Bardella’s remarks “aim to deflect the scrutiny” amid an approaching election, attributing a motivation that wasn’t part of the source material.
As to whether AI has any place in journalism, Kaeberlein and Alikhani agreed — it’s context-dependent and should be used carefully.
“It’s important to keep and protect human judgement, human agency,” Alikhani said. “Because we know AI can be biased, we know AI can hallucinate, we know AI can suggest a different tone that we may not catch, evaluate, or shift after we read it.”
Their findings raised questions about how to make the technology more accountable. Looking ahead, they hope to research transparency measures and LLMs’ ability to evaluate themselves.
“As much as we want AI to be human-like, for us to be able to interpret it, for us to be able to work with it, they’re not people,” Alikhani said. “Don’t outsource ethics.”
The Khoury Network: Be in the know
Subscribe now to our monthly newsletter for the latest stories and achievements of our students and faculty