Experimenting with Perplexity’s Deep Research: Real Sources, Fake Facts
A sobering reminder we can't trust the the model.
I recently shared my favorite ways to leverage AI in People Research & Analytics, but the AI landscape is constantly evolving and introducing new, exciting features. Last month, Perplexity launched its new “Deep Research” algorithm, which promises “to save you hours of time by conducting in-depth research and analysis on your behalf.” As a researcher, this was music to my ears. So I dove in.
Why was I excited to try Deep Research?
It’s free
Perplexity has a better reputation—at least comparatively—for generating real (not hallucinated) sources
Ezra Klein recently said on his podcast that Deep Research was able to generate a report in minutes that was on par with the “median” of what his team of researchers would normally produce in days (a pretty strong endorsement, in my opinion)
The Experiment
To test out the feature, I decided to start with a very “meta” question that has been on my mind lately: how AI, and specifically Artificial General Intelligence (AGI), might impact my role in People Analytics long term:
“Researchers speculate that AGI is on the horizon in the next few years. Knowledge workers like myself are particularly at risk. I work as a people analyst/people research scientist. On the one hand, AI and AGI can help me do my job more efficiently, but on the other hand, it might make me replaceable. How at risk do you think my role is of becoming obsolete and what are skills I can develop or actions I can take that will ensure I am able to add value above and beyond AGI in the future (including skills related to leveraging AI/AGI)?”
The model took a couple minutes to run, keeping me apprised of the steps it was doing along the way. It then produced an answer of just over 1,000 words.
But was it any good?
What the model did well:
Real sources: all of the 9 cited sources linked to real sources (not hallucinated citations).
Content structure: The response was well-organized, featuring an intro, clear sections addressing different aspects of my question, and a conclusion. I do think the narrative and cohesion of the content could have been tighter, but not bad for a 2 minute wait.
Some insights: The response included some novel suggestions for how I could leverage AI in my role and skills I may want to develop.
What the model did poorly:
Quality of the sources: Of the 9 sources cited in the answer, none were from academic journals. Only one was what I would consider an “official” news sources (CBS News); the others were personal substacks (no shade, we love those), small tech blogs, glorified marketing content (e.g., articles written on the website of a company selling their product), and one was just the landing page for a People Analytics Summit.
Fake facts and statistics: Although the sources were real, the model was abysmal at accurately conveying the findings and statistics from those sources. More on this below.
With regard to the quality of the sources, perhaps my academic bias is showing: when citing facts and findings, I would strongly prefer a peer reviewed paper over a sponsored article or random tech blog. That said, for some purposes those alternative sources might be sufficient.
However, what I cannot get past is the blatant disconnect between the summarized information and the actual content of the sources. Here are just a few examples of what Deep Research told me:
“Current generative AI systems demonstrate 55-70% proficiency in core people analytics functions according to industry benchmarks.” The cited source (a marketing page for an HR software company) makes NO mention of this 55-70% statistic. At best, it qualitatively describes some of the ways that AI can help automate HR workflows. In fact, the only percentage referenced in the source is a completely different one (“42% of firms with at least 5,000 workers reported using AI for HR tasks in early 2022”), which Deep Research failed to mention in its summary.
“AI systems now process 92% of exit interview analysis and 78% of promotion pipeline forecasting at Meta, according to leaked 2024 implementation reports.” As someone who works in People Research at Meta, I took this one personally. The cited source–a Palo Alto based tech-blog–discusses how Meta may introduce AI-generated content into its public-facing apps. It had absolutely nothing to do with hiring, promotions, or HR processes. That’s quite a leap, Deep Research.
“A 2024 Stanford study found that human-AI teams outperformed pure AI systems by 41% in predicting leadership succession conflicts within tech firms.” This one might be my favorite because it had such promise and is perhaps the most egregious. The actual linked source is an article on myhrfuture.com called “Five Core Skills for People Analytics.” There is no Stanford study. No mention of leadership succession or conflicts in firms. Nothing.
I counted a total of 10 statistics—hard numbers—cited in Deep Research’s answer to my question. Of those, all ten were unsupported by the cited sources.
Ironically, one of the sources DID include a real statistic relevant to my question: a report from the job website Indeed stating that 20% of roles were highly exposed to replacement by AI, 45% moderately exposed, and 35% minimally exposed. Yet, bafflingly, Deep Research failed to include that.
But everybody deserves a second chance, right?
I thought perhaps the issue was that I had asked Perplexity to provide insight on a topic where the extant research was limited, so it resorted to making things up. So I tried again, this time asking for research on a topic that has been extensively researched and reported on: leadership.
“I want to develop a model of leader effectiveness based on scientific search. While leaders are also managers, leadership is arguably broader than just management. What would be a concise, research-backed model of leadership effectiveness that, critically, could actually be measured, tracked, and cultivated in an organization?”
What the model did well:
Real sources: Once again, all of the 4 cited sources linked to real articles.
Some insights: And once again, the answer conveyed some legitimate themes in the research on leadership and leader effectiveness.
What the model did poorly:
Quality of the sources: In the model’s defense, this time it did cite what appears to be a peer-reviewed paper; however, the remaining 3 sources were once again a collection of marketing content and LinkedIn posts.
Fake facts and statistics: My hopes were once again dashed. Right out of the gate, Deep Research offered this gem: “Studies correlate high cognitive flexibility with 23% faster decision-making in crisis scenarios.” This was in no way supported by the cited source. In fact, when I searched the Internet, I couldn’t find any record of this statistic.
I won’t subject you to another step-by-step walkthrough of the hallucinated facts, but it was a repeat of my first test. And while I’ve seen human-written literature reviews engage in some generous extrapolation when citing sources, Deep Research took it to a new level.
The Takeaway
Look, I’m excited about the potential of tools like Deep Research. I wanted it to blow me away, and it may well do so in the future. Ironically, AI is a useful tool for detecting errors in scientific research, but as my experiment with Perplexity’s new model shows, it can also introduce them.
So where does this leave us? As we adopt AI, we must remain vigilant. We risk becoming complacent by believing that the systems are reliable, that they will protect us from mistakes, and if the system was doing something seriously harmful, someone would step in. Only no one is coming, not yet. You have to be that someone.
YOU are the guardrail when using AI output.
So experiment. Use AI. But verify everything. Because in a world where AI confidently cites fake facts from real sources, skepticism isn’t just useful—it’s necessary.

