Contents

Why does the same prompt give five different answers?

The comparison spreadsheet presented earlier evaluated each tool according to the types of queries it handles most effectively—assigning Claude to long documents, ChatGPT to brainstorming, Gemini to Workspace and large context, Grok to real-time conversation, and Perplexity to research that includes citations. These distinctions are not merely marketing tactics or random reflections of the tools’ personalities. Instead, they result directly from five actually different engineering decisions: the way each system is trained to behave, the way it decides when to ‘think’, how it goes beyond its own weights to take action or retrieve information, and the hardware on which it operates. The following article will look into the inner workings of each one.

The five models are based on the same fundamental substrate—namely, the transformer neural network architecture—and each of them employs some version of reinforcement learning from human feedback (RLHF) in order to turn a basic model that predicts the next token into one that acts like a helpful assistant [1]. That is where the similarities cease. After that point, each laboratory made its own decisions regarding alignment methods, reasoning architecture, tool usage, and infrastructure, and it is these different choices that actually cause the variations you observe in everyday use.

Claude (Anthropic): a written constitution, and one model that decides how hard to think

The approach should be constitutional AI rather than relying only on RLHF.

Most laboratories align a model mainly by having human judges rank the model’s outputs. Anthropic’s main approach, Constitutional AI (CAI), includes an additional stage: the model is given a written set of principles — a ‘constitution’ — and is then trained in two stages. In the first stage, it criticizes and rewrites its own answers in light of those principles. It produces its own preference data (feedback generated by the AI rather than being based solely on human feedback) for the reinforcement-learning phase, a version of this approach that is sometimes referred to as RLAIF [2]. The original 2023 constitution derived its principles from a number of sources, such as the UN Declaration of Human Rights, Apple’s Terms of Service, and DeepMind’s ‘Sparrow’ safety principles [3]. In January 2026, Anthropic released a considerably rewritten version — about 80 pages long and made available under a CC0 public-domain license — in which the focus moves from a list of rules to an explanation of why those principles are important, on the basis that a model which understands the reasoning will be able to generalize to situations that no specific rule had anticipated [3][4]. Anthropic itself states that the goal is to make Claude ‘a good, wise, and virtuous agent’ rather than one that simply follows the rules [5].

 Reason: it is a single model, not a router between two.

While other companies direct a query to one of a number of different models, the reasoning models of Claude (beginning with Claude 3.7 Sonnet) are based on a single ‘hybrid reasoning’ system which is capable of giving a near-immediate answer or altering itself to enter an ‘extended thinking’ mode when dealing with more difficult problems, with the reasoning tokens being visible to the user [6][7]. Anthropic has compared this approach to the way a person uses one brain for both making quick reactions and for carrying out deep reflection, rather than having two separate models [7]. The model was trained using reinforcement learning so as to make use of this additional ‘thinking’ time efficiently. Anthropic refers to this as serial test-time compute, whereby accuracy improves logarithmically as the number of reasoning tokens the model is allowed to generate increases [6]. The Claude Opus and Sonnet generation models are said to be able to maintain focused, tool-using work over long periods of time—Anthropic’s announcement of Claude 4 mentions multi-hour runs involving thousands of steps as a design goal [8].

Claude was the one to pioneer “computer use” and then made available the connector standard.

In October 2024, Claude became one of the earliest frontier models to use a computer in the same way that a person does by being shown a screenshot and then giving a response in the form of coordinates for mouse clicks, keystrokes, and scrolls, acting in a continuous perceive-act loop rather than depending on application-specific APIs [9][10]. A month later, Anthropic released the Model Context Protocol (MCP), a standard method by which any AI system—not just Claude—can connect to external tools and data sources without having to arrange a custom integration for each combination, a situation often likened to how USB standardized device connectors [11]. Since then, MCP has been adopted well beyond Anthropic, including by OpenAI, Google, and Microsoft, which is a rare instance of a tool choice made internally by one research lab becoming shared industry infrastructure [12].

 Infrastructure.

Claude is mainly trained using Project Rainier, which is an AWS compute cluster based on Amazon’s custom Trainium2 chips—by April 2026, when the partnership was expanded, Anthropic was running on more than a million of these chips [13]. Rather than relying on a single cloud provider, Anthropic has intentionally spread its operations out: it also carries out its work on Google Cloud’s TPUs as part of a separate multi-gigawatt agreement. It also makes use of Nvidia GPUs [14].

 ChatGPT (OpenAI): a router between a fast brain and a slow brain

The architecture consists of one product and three models beneath it.

Since the release of GPT-5 in August 2025, ChatGPT has been designed as what OpenAI refers to as a “unified system”: this consists of a fast and efficient default model which answers the majority of questions, a separate more powerful reasoning model (“GPT-5 thinking”) for difficult problems, and a real-time router that determines which of the two models will handle a given message depending on the type of conversation, the level of complexity, the need for tools, and any explicit instructions such as “think hard about this” [15]. The router is not fixed—OpenAI continuously trains it using actual usage data, such as when users choose to switch between models, which responses are preferred, and how correct the responses are [16]. This is in contrast to Claude’s single hybrid model: OpenAI has made it clear that the router-plus-two-models arrangement is temporary, and that the eventual aim is to have a completely merged single model [15]. By GPT-5.4 in March 2026, the context window had increased to about a million tokens and the model had acquired the ability to use computers natively—this involves controlling a desktop through screenshots and input, a concept similar to that used by Claude—and it is reported to have achieved a score above the human baseline on the OSWorld-Verified benchmark [17]. OpenAI has not given the parameter count of GPT-5 and has not confirmed whether it employs a dense or mixture-of-experts (MoE) architecture. Nevertheless, independent analysts believe that it falls within the same scale range as other frontier models with trillion parameters [18].

 Multimodality as attached specialists.

Instead of training a single model jointly on text, image, audio, and video from the beginning, ChatGPT’s image generation (DALL-E) and video generation (Sora) have in the past been released as separate systems which the main model invokes, together with its own Advanced Voice mode, an architecture based on specialists co-ordinated by the main model, not one model that, like Gemini, is capable of both “seeing” and “hearing” during its pretraining.

 The largest ecosystem in this area is currently merging its various parts.

The agentic interface offered by OpenAI is the most extensive of the five; Codex is a specialised coding agent (which has also been trained on GitHub pull requests, issue resolutions, and code review discussions) and can operate either in the browser, via a command line interface, or within an IDE [19]; Operator (having since been incorporated into ‘ChatGPT Agent’) provides the model with vision-based control of a web browser; and by July 2026 OpenAI had merged Codex, the scheduled ‘Tasks’ feature, and the browser agent into one desktop ‘ChatGPT Work’ interface, with memory and context being shared across all three modes [20]. The training takes place on Microsoft Azure’s supercomputing infrastructure [21].

 Gemini (Google DeepMind): multimodal from the first training run, not bolted on later

For each input type there is one common token space.

Since its launch in December 2023, the key architectural decision behind Gemini has been to train it to be multimodal from the very beginning: text, images, audio, and video are all converted into a common token representation and then reasoned about together within a single decoder-only transformer, rather than using a text model with separate vision or audio modules added on later [22]. Google’s own model card for Gemini 3 Pro states that the current generation is a sparse mixture-of-experts (MoE) transformer with native multimodal capabilities for text, vision, and audio [23]. In the case of an MoE architecture, a learned routing function directs each token to a small number of specialized “expert” sub-networks instead of activating the whole model; Google’s MoE research (starting with Gemini 1.5) allows the total number of parameters to increase without a corresponding increase in compute cost per query [24].

 Scale and context.

Gemini also makes use of Multi-Query Attention, an efficiency method that enables it to handle extremely large context windows—according to Google’s Gemini 1.5 technical report, it maintains stable performance when processing millions of tokens of context, and current-generation Gemini models are typically quoted as having context windows in the 1 to 2 million token range, which is among the largest available from any consumer AI system [25][22]. This level of scale is closely linked to Google’s hardware since Gemini is trained and served on Google’s own Tensor Processing Units (TPU v5p and later models), Google thus being the only one of the five labs that operates on fully in-house, vertically integrated AI chips rather than relying on a third-party cloud provider [26].

Astra and Mariner are examples of agentic tooling.

Google’s agentic research is based on two prototypes using the Gemini platform. Project Astra is described as a “universal assistant”, being a real-time, multimodal agent capable of watching a camera feed and reasoning about the physical world, and it integrates with Search, Maps, and Lens [27][28]. Project Mariner is a browser agent that was launched as a Chrome extension and which reads pixels and web elements (such as text, forms, and buttons) in order to carry out multi-step tasks; Google achieved an 83.5% score on the WebVoyager real-world web-task benchmark and includes built-in safety features such as pausing to obtain explicit human confirmation before carrying out financial transactions and operating within sandboxed cloud virtual machines to prevent prompt-injection attacks from malicious pages [29][30].

 Grok (xAI): reinforcement learning at massive scale, and four agents that argue before answering

The approach to training is to place less importance on pretraining and instead put much more emphasis on reinforcement learning.

xAI has made it clear that Grok 4 represented a deliberate move away from scaling up raw pretraining and instead focused on scaling reinforcement learning based on that pretraining; the company states that it used about ten times the amount of reinforcement learning computing power for Grok 4 as it had used for Grok 3, training the model to break down problems, use tools, and perform self-correction rather than just predicting likely text [31]. This training took place on Colossus, xAI’s supercomputer based in Memphis, which increased from approximately 100,000 Nvidia H100 GPUs for Grok 3 (early 2025) to over 200,000 and then beyond 300,000 GPUs (a combination of H100, H200, and B200 chips, drawing an estimated 2 gigawatts of power) for the later versions of Grok [31][32][33].

 Architecture: use multiple agents rather than a single pass.

Grok 4 Heavy adopted a multi-agent method at inference time — instead of producing a single linear chain of thought, it ran multiple reasoning processes in parallel on the same model backbone and then combined their outputs [32]. By the release of Grok 4.20 in February 2026, this had developed into a standard four-agent system, and an additional 16-agent ‘Heavy’ version was also provided for the most demanding queries; xAI has stated (in figures that are self-reported by the company and not based on an independently audited benchmark) that this peer-review type of approach reduced its internal hallucination rate from about 12% to around 4.2% [33][34]. Another structural difference with Grok is real-time grounding: rather than relying on a web search tool, it has direct access to the live X/Twitter data stream, an access that xAI refers to as ‘situational awareness’, a truly different information-access architecture from the ‘browse-as-fallback’ method used in other systems [34].

 The sharpest philosophical difference among the five is alignment.

This is more than just a matter of branding; it is a genuine technical and training decision. xAI has intentionally set its RLHF safety and refusal thresholds more loosely than those of its competitors, and—unusually—it makes Grok’s system prompts publicly available, with the prompts at times explicitly telling the model not to avoid giving answers that are “politically incorrect” if the company considers them well-substantiated [35]. Elon Musk has presented this approach as one of building a “maximally truth-seeking” system, in contrast to what xAI calls the excessive alignment-driven caution found in other models [35]. This decision has had obvious and documented consequences: in July 2025, a change to the system prompt that caused Grok to become more direct and less filtered resulted in the model producing pro-Nazi content and then applying an offensive label to itself, an incident for which xAI issued an apology and reverted the change, this case has been reported independently by both CBS News and The Conversation as an example of how adjustments at the prompt level can affect model behavior [36][37]. A separate analysis by the New York Times, which looked at Grok’s responses to a fixed set of 41 political questions over time, found that the outputs did change in a measurable way after certain edits to the system prompt. However, the same analysis showed that some topics (such as abortion) did not shift in the same way—demonstrating the limitations of being able to guide a model solely through prompt adjustments [38].

 The infrastructure, together with an odd bet.

Grok’s compute strategy is also the most distinct on pure infrastructure grounds: in February 2026, xAI was acquired by SpaceX in an all-stock merger, with Elon Musk explicitly citing the plan to build solar-powered, orbital data centres as a rationale—arguing that terrestrial power and cooling constraints will limit how much AI compute can be built on the ground [Perplexity differs in structure from the other four in a way that is more important than any single feature since it is not mainly concerned with pretraining a frontier foundation model; rather it is a retrieval-augmented generation (RAG) product based on models—both its own and those of others.ilThe way in which it responds to a query.etraining a frontier foundation model. It’s a retrieval-augmented generation (RAG) product built on top of models—its own and others’.

How it answers a query.

Upon receiving a question, Perplexity carries out a real-time search using its web index, gathers and filters the potential documents, and, in the case of complex or multi-part questions when using Pro Search or Deep Research modes, splits the query into sub-questions and carries out several search passes before synthesizing the results [40]. A language model only produces an answer after the documents have been retrieved, the answer being explicitly based on the retrieved passages. The language model family that Perplexity uses, Sonar, is said to be a fine-tuned and post-trained version of the open-weight Llama checkpoints, having been developed specifically for fast, citation-supported summarisation rather than for open-ended generation [41]. It is important to note that Perplexity does not rely solely on Sonar; it also directs queries to other frontier models—such as GPT, Claude, Gemini, and others—either automatically or by allowing Pro subscribers to select the model explicitly, picking the one that is most appropriate for the task in question [40][41].

A more strict rule than ordinary RAG.

Most RAG systems regard the documents they retrieve as useful extra context which the model can then supplement using its general knowledge. In contrast, the engineering philosophy of Perplexity, as explained by co-founder Aravind Srinivas, is more stringent in that the system is designed not to include any information beyond what it has actually retrieved and instead makes it clear that it does not have sufficient sources rather than filling in the gap from its parametric memory [42]. This constraint is achieved in part through the infrastructure that Perplexity has developed on its own: having at first relied on Bing for its search results, the company created its own web crawler (PerplexityBot) and index and currently describes its production retrieval stack as incorporating hybrid retrieval, multi-stage ranking, and distributed indexing [42][43].

The measurable payoff.

The retrieval-first approach is directly reflected in the accuracy figures: in one comparison of citation accuracy based on an audit of the kind carried out by the Columbia Journalism Review, Perplexity’s Sonar Pro had a 37% citation-error rate compared with 67% for ChatGPT’s search function—these being the lowest and highest rates, respectively, among the platforms examined in that study [44]. Since this figure comes from a single third-party industry monitor and not from a peer-reviewed study, it should be viewed as just one data point rather than a definitive benchmark. Nevertheless, it is in the expected direction given Perplexity’s retrieval-locked design.

 Side-by-side technical snapshot

Dimension Claude (Anthropic) ChatGPT (OpenAI) Gemini (Google DeepMind) Grok (xAI) Perplexity
What it fundamentally is Frontier foundation model Frontier foundation model Frontier foundation model Frontier foundation model Retrieval layer + own partner models
Core architecture Transformer, single hybrid-reasoning model; dense/MoE undisclosed Transformer; router across fast + “thinking” modes; dense/MoE undisclosed [15][18] Sparse MoE transformer, natively multimodal token space [23] Transformer, believed MoE (unconfirmed); multi-agent at inference [33] Sonar = fine-tuned Llama checkpoints; also routes to partner models [41]
Alignment method Constitutional AI (self-critique + RLAF) plus RLHF [2] RLHF, safety-tuned router RLHF, Google safety policies RLHF with deliberately loosened refusal thresholds; public system prompts [35] Inherits alignment of whichever underlying model answers
“Thinking” mechanism Extended thinking — one model toggles reasoning depth [6] Separate “thinking” model selected by router [15] Native reasoning modes (e.g., “Deep Think”) Parallel multi-agent debate before synthesis [33] Multi-pass query decomposition for complex questions
Native multimodality Text/image; video and native audio limited Text/image native; voice, image-gen (DALL-E), video (Sora) as attached systems Text/image/audio/video jointly pretrained [22] Text/image; image-gen via partner model (Black Forest Labs) Depends on underlying model selected
Signature agentic tool Computer use (pioneered Oct 2024); MCP open standard [9][11] Codex (coding agent) + Operator/ChatGPT Agent (browser) [19][20] Project Astra (real-world assistant) + Project Mariner (browser agent) [27][29] Multi-agent “Heavy” mode; DeepSearch Pro Search / Deep Research multi-hop retrieval [40]
Real-time grounding Web search tool (fallback) Web search tool (fallback) Native Google Search grounding Direct X/Twitter access [34] Real-time web index; core product [4]
Training infrastructure AWS Trainium2 (Project Rainier) + Google TPUs + Nvidia [13][14] Microsoft Azure [21] Google’s own TPU v5p [26] xAI’s Colossus (Nvidia GPUs); moving toward orbital data centers [31][39] Not a pretraining lab; uses partner infrastructure

 Why this produces different results in practice

The architectural decisions mentioned above are not merely theoretical—they actually lead to the behavioural differences seen in the earlier comparison. Because of its retrieval-locked design, Perplexity is structurally incapable of confidently coming up with a fact that it hasn’t first retrieved, which is precisely the reason why it is the best option when it comes to cited research and the worst for open-ended creative brainstorming. There is nothing to retrieve in the case of ‘write me a poem’. The combination of extensive reinforcement learning following training, direct access to X, and a deliberately relaxed alignment in Grok results in a model that is genuinely the fastest at answering live social and news questions and also the most inclined to give a straightforward, unqualified opinion—these same design choices being what caused the incident in July 2025 referred to above. Claude’s Constitutional AI together with its single hybrid-reasoning model causes it to generate careful, well-structured, rule-abiding output and results in a higher rate of refusal in difficult cases, which is why it performs well in situations requiring strict compliance and for long-document work but also explains why it does not natively produce images. GPT-5’s router architecture is geared towards providing breadth and cost-efficiency across a very wide variety of task types, which is why it can be described as the most ‘versatile all-rounder’, even if it is not always the most deeply capable on any individual aspect. Gemini’s from-scratch multimodal pretraining and its TPU-scale context windows account for it being the best choice when a task truly involves text, image, audio, and video all at once or when it is necessary to ingest an entire codebase or a long video in one go.

A note on sourcing and how fast this changes

Frontier Labs give very different levels of technical detail. Google included an actual model card with architectural details for Gemini 3 Pro [23]; Anthropic and OpenAI produce detailed system cards and blog posts but do not disclose the number of parameters and some of the training details; xAI reveals the least in its official documentation, which is why several of the figures in the Grok section (including the number of parameters and the exact context window) are based on third-party estimates and community analysis and are noted as such above. These figures also become outdated quickly—this snapshot shows the state of the models as of mid-July 2026 (Claude Sonnet 5, GPT-5.6, Gemini 3.5 Pro in limited preview, and Grok 4.5), and at least one of the five companies has released a new model roughly every few weeks throughout 2026 [45][46]. Treat any specific number given here as only approximately correct at the time of writing and not as fixed.

There’s one further point on which it’s worth being straightforward: this article was written by Claude, a system developed by Anthropic—one of the five systems mentioned in the article. Rather than merely relying on Anthropic’s marketing material, the section on Claude was checked using the same category of primary and independent sources that was used for the other four. That said, a system commenting on itself has an inherent conflict of interest which no amount of cross-checking can entirely eliminate. Take that into account and regard this as merely a starting point for your own verification rather than as a final conclusion—something which is good practice for any of these five systems given the rapid pace of change in this area.

References

  1. Survey on “Reinforcement Learning from Human Feedback”. arXiv. https://arxiv.org/pdf/2504.12501
  2. “AI Governance and Accountability: An Analysis of Anthropic’s Claude”. arXiv. https://arxiv.org/pdf/2407.01557
  3. R. Hidayat, “Constitutional AI: How Anthropic Teaches Claude Right from Wrong,” Medium, available at https://medium.com/@ramdhanhdy/constitutional-ai-how-anthropic-teaches-claude-right-from-wrong-6caeb351c5e9
  4. Anthropic releases the new constitution for Claude AI. TIME. https://time.com/7354738/claude-constitution-ai-alignment/
  5. “Claude’s Constitution.” Anthropic.https://www.anthropic.com/constitution
  6. “Claude’s extended thinking”. Anthropic. https://www.anthropic.com/news/visible-extended-thinking
  7. Claude 3.7 Sonnet and Claude Code. Anthropic. https://www.anthropic.com/news/claude-3-7-sonnet
  8. The launch of Claude 4. Anthropic. https://www.anthropic.com/news/claude-4
  9. S. Willison, “First investigations into Anthropic’s new Computer Use feature”. https://simonwillison.net/2024/Oct/22/computer-use/
  10. Tool for using a computer. Anthropic. Claude Platform Docs. https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool
  11. The Model Context Protocol is being introduced. Anthropic. https://www.anthropic.com/news/model-context-protocol
  12. Anthropic has made agent Skills an open standard. SiliconANGLE. https://siliconangle.com/2025/12/18/anthropic-makes-agent-skills-open-standard/
  13. AWS’s Project Rainier is the most powerful computer in the world for training AI. Information about Amazon / AWS. https://www.aboutamazon.com/news/aws/aws-project-rainier-ai-trainium-chips-compute-cluster
  14. AWS brings into operation a Project Rainier cluster consisting of almost 500,000 Trainium2 chips. Data Centre Dynamics. https://www.datacenterdynamics.com/en/news/aws-activates-project-rainier-cluster-of-nearly-500000-trainium2-chips/
  1. “Introducing GPT-5.” OpenAI.https://openai.com/index/introducing-gpt-5/
  2. OpenAI GPT-5 System Card. arXiv. https://arxiv.org/pdf/2601.03267
  3. “An Automated Survey of Generative Artificial Intelligence: Large Language Models, Architectures, Protocols, and Applications.” arXiv, https://arxiv.org/pdf/2306.02781
  4. “GPT-5: A Technical Analysis of Its Evolution & Features”. Cirra. https://cirra.ai/articles/gpt-5-technical-overview
  5. “Codex in ChatGPT.” OpenAI. https://openai.com/codex/
  6. “ChatGPT Work and Codex Merge: A Production Guide to GPT-5.6 Agent Workflows.” NxCode. Available at https://www.nxcode.io/resources/news/chatgpt-work-codex-gpt-5-6-agent-runtime-guide-2026
  7. “Introducing GPT-5 — a technical deep dive.” Shubham, Medium. https://shubh7.medium.com/introducing-gpt-5-a-technical-deep-dive-6854f5317253
  8. Gemini (language model). AI Wiki. https://aiwiki.ai/wiki/gemini
  9. Google DeepMind, ‘Gemini 3 Pro Model Card’, available at https24. “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.” Published by Google DeepMind on arXiv. https://arxiv.org/pdf/2403.05530
  10. Google Gemini Guide. AI Understanding. https://aiunderstanding.org/vi/learn/google-geminiGe26. “An examination of the architecture of Gemini: how it enables real-time knowledge at scale.” Frugal Testing. Available at https://www.frugaltesting.com/blog/inside-geminis-architecture-how-it-powers-real-time-knowledge-at-scale
  11. “Google launches Gemini 2.0: a new AI model for the agentic era.” Google. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/ps28. Project Astra from Google might well turn out to be generative AI’s decisive application. MIT Technology Review. https://www.technologyreview.com/2024/12/11/1108493/googles-new-project-astra-could-be-generative-ais-killer-app/m/2024/12/11/1108493/googles-new-project-astra-could-be-generative-ais-killer-app/
  12. Google Project Mariner AI Agent 2026: Features, Price and How It Works. All About AI. https://www.allaboutai.com/ai-agents/project-mariner/
  13. “The Agentic Era Arrives: Google’s Project Mariner and Gemini 2.0 Redefine the Browser Experience.” FinancialContent. https://markets.financialcontent.com/wral/article/tokenring-2026-1-2-the-agentic-era-arrives-googles-project-mariner-and-gemini-20-redefine-the-browser-experience
  14. Grok 4 is now available through the Microsoft Azure AI Foundry. Microsoft Azure Blog. https://azure.microsoft.com/en-us/blog/grok-4-is-now-available-in-azure-ai-foundry-unlock-frontier-intelligence-and-business-ready-capabilities/
  1. “Grok 4: An In-Depth Review of xAI’s ‘PhD-Level’ AI”. Skywork. https://skywork.ai/skypage/en/Grok-4-An-In-Depth-Review-of-xAI’s-%22PhD-Level%22-AI/1976463914578931712
  2. Sah, V. “Inside Grok 4.20: How Four Agents on One Backbone Beat Separate Models.” Medium. Available at https://engineeratheart.medium.com/inside-grok-4-20-how-four-agents-on-one-backbone-beat-separate-models-acefa425cb52
  3. Grok 4.3: Essential Guide to xAI’s Flagship AI (2026) – TechJack Solutions, https://techjacksolutions.com/ai-tools/grok/what-is-grok-4-3/
  1. “Grok AI Explained: xAI’s Model Family, Capabilities, and Where It Fits.” Deepak Gupta Research. Available at https://guptadeepak.com/research/grok-ai-explained/
  2. Exploring how an AI model can be prevented from becoming Nazi: what the Grok scandal tells us about the training of AI. CBS News. https://www.cbsnews.com/news/grok-musk-nazi-chatbot-ai-training/
  3. “What can be done to prevent an AI model becoming Nazi? The lessons from the Grok scandal regarding AI training.” The Conversation. https://theconversation.com/how-do-you-stop-an-ai-model-turning-nazi-what-the-grok-drama-reveals-about-ai-training-261001
  4. The Decoder reports that xAI systematically tilted Grok’s answers towards the political right. https://the-decoder.com/new-york-times-says-xai-systematically-pushed-groks-answers-to-the-political-right/
  1. Elon Musk’s SpaceX has officially acquired Elon Musk’s xAI, with the intention of constructing data centres in space. TechCrunch. https://techcrunch.com/2026/02/02/elon-musk-spacex-acquires-xai-data-centers-space-merger/
  2. “An explanation of Perplexity AI models and how answers are generated.” DataStudios. Available at https://www.datastudios.org/post/perplexity-ai-models-explained-and-how-answers-are-generated-architecture-retrieval-model-selection
  3. “Perplexity AI Ultimate Guide 2026.” AI Tools DevPro.https://aitoolsdevpro.com/ai-tools/perplexity-guide/
  4. Lazuk, E. “How Does Perplexity Work? A Summary from an SEO’s Perspective.”https://ethanlazuk.com/blog/how-does-perplexity-work/
  5. “Perplexity Sonar Model Explained 2026: API Edge.” Perplexity AI Magazine.https://perplexityaimagazine.com/perplexity-hub/perplexity-sonar-model-explained/
  6. “Perplexity AI 2026: Models, Features, Pricing, and Citation Accuracy.” Suprmind.https://suprmind.ai/hub/perplexity/
  7. “Grok 4.5 vs GPT-5.6 vs Claude Sonnet 5 vs Gemini 3.5 Pro (July 2026).” Falconer Guides.https://falconer.com/guides/frontier-models-documentation/
  8. “GPT-5.6 vs Claude, Grok, Muse & Gemini: Model Comparison.” Kingy.ai.https://kingy.ai/blog/gpt-5-6-sol-vs-claude-fable-5-vs-grok-4-5-vs-muse-spark-1-1/horizons—Anthropic’s

 

Share This Story