narendra raghunath
a story illustration of Anekantavada
(I decided to step away from my practice-based PhD research after an initial period of engagement, not because I lost interest in practice-based inquiry, but because the experience led me to question how knowledge is constructed, organised and validated within higher education in India. I encountered a continuing tension between the aspiration to make academic practice more inclusive and postcolonial and the persistence of institutional conventions that privilege particular ways of defining a problem, establishing evidence, constructing an argument and presenting a conclusion. English remains the dominant language of advanced academic communication in many Indian universities, and with it comes not only a language structuring but also established conventions of scholarly expression and validation. My concern is not that these conventions are inherently Western or inherently inadequate, but that they can become so institutionalised that they are treated as neutral standards even when they do not fully accommodate knowledge produced through artistic practice, embodied experience, iterative making or relationships among different intellectual traditions.
This experience made me reconsider what constitutes accuracy in scholarship. Every knowledge system must distinguish fact from interpretation, analogy from influence, metaphor from evidence, correlation from causation, perspective from contradiction, and synthesis from integration. But another distinction is equally important: the difference between the accuracy of an individual claim and the accuracy of the relationships through which claims acquire meaning. I describe the latter here as relational or structural fidelity. Factual accuracy asks whether a particular statement is supported by reliable evidence; while structural fidelity asks whether the relationships among claims, concepts, sources, contexts and perspectives have been preserved. An argument may therefore be factually correct while remaining historically, philosophically or conceptually misleading.
For example, two artistic traditions may look similar without one influencing the other. Events may occur one after another without one causing the other; a translated philosophical term may not carry the full meaning of the original concept; and a clear historical narrative may still simplify or distort a disputed past.
The meaning of an idea often depends on its relationship with other ideas, the questions it addresses, the assumptions it rejects, and the historical or material practices through which it gains meaning. This distinction became important to me because knowledge is rarely made up of isolated propositions. When these relationships are removed, a statement may remain correct while its intellectual meaning changes. My experience with practice-based research made me question whether academic systems sometimes value the form in which knowledge is presented more than the conditions through which it is produced.
From this emerged a second question: if academic systems can regularise diverse forms of knowledge into particular structures of validation, what happens when artificial intelligence learns from and reproduces the knowledge contained within those systems? Could an LLM inherit not only information from its training materials but also some of the tendencies through which information is selected, framed, connected and interpreted?
This question became more concrete when I read the 2024 study by Yan Tao, Olga Viberg, Ryan S. Baker, and René F. Kizilcec, published in the paper "Cultural Bias and Cultural Alignment of Large Language Models," in PNAS Nexus. The researchers compared five GPT models with nationally representative cultural-value data covering 107 countries and territories and found systematic differences between model outputs and local cultural values. In their evaluation, the tested models' default outputs were closer to cultural-value patterns observed in English-speaking and Protestant European countries, whereas culturally specific prompts improved alignment across many countries. The study does not show that LLMs follow a single cultural or argumentative structure, but it does provide evidence that general-purpose models exhibit systematic cultural patterns in their outputs.
For me, this opened a broader line of inquiry. The question was no longer simply whether an academic institution is prepared to recognise forms of knowledge that do not fit comfortably within its established structures. It became possible to ask whether the increasingly influential computational systems through which knowledge is now searched, summarised and reorganised may reproduce some of the same structural tendencies in another form.
This does not mean that academic systems and LLMs are identical, nor that LLMs reproduce one particular intellectual tradition. Their behaviour is shaped by training data, model design, post-training, prompting, system instructions and other computational processes. The concern is more specific: when diverse bodies of knowledge are repeatedly reorganised into highly coherent linguistic forms, relationships essential to their original meaning may be weakened, simplified, or reframed.
The following essay, therefore, examines AI-mediated knowledge from this point of intersection. It asks whether factual correctness alone is enough for scholarly reliability, or whether responsible knowledge production must also preserve the conceptual, historical and relational structures through which ideas acquire meaning. It also considers what may happen to intellectual plurality when systems capable of producing and circulating arguments on an enormous scale increasingly shape how those arguments are constructed, presented and understood.
The central concern of the essay is thus not whether AI can replace scholarship, but whether the architecture through which AI organises knowledge can itself become an unexamined part of scholarship.)
Essay :
Large language models have changed the relationship between information and writing by enabling the retrieval, comparison, summarisation, and reorganisation of enormous quantities of textual material within seconds. Their usefulness in scholarship, however, raises a question that is deeper than the familiar problem of factual accuracy. An answer can contain correct facts, accurate quotations and legitimate references and still produce a misleading argument because the relationships among those facts have been reorganised in a way that changes their meaning. This is particularly important when an inquiry brings together philosophical, historical, scientific, artistic, or culturally situated forms of knowledge, because intellectual traditions do not always organise propositions, evidence, contradictions, perspectives, and contexts in the same way. The central problem is therefore not that large language models necessarily follow one particular intellectual tradition, but that they can transform heterogeneous forms of reasoning into comparatively regular, coherent and familiar linguistic structures. The result may be a form of structural simplification in which the individual components of an argument remain recognisable while the relationships that originally connected them are weakened or altered.
This argument needs to begin with a qualification. It would be historically and technically inaccurate to say that LLMs reproduce a Western model of thought. Western thought is a much larger history of European reasoning, and European intellectual traditions include dialectical, phenomenological, hermeneutic, pragmatic, systems-oriented and other approaches that cannot be reduced to syllogistic reasoning. At the same time, many intellectual traditions outside Europe developed highly formal systems of inference, debate and epistemology. Classical Indian traditions, for example, include extensive discussions of inference, valid cognition, testimony, contradiction and the conditions under which knowledge can be established. The important point is therefore not that one civilisation possesses "linear logic" while another possesses "holistic thinking". Such a division would reproduce precisely the kind of simplification that this essay seeks to question. The more defensible proposition is that different intellectual traditions organise relationships among propositions, perspectives and contexts in different ways, and that an LLM can sometimes flatten those differences when it converts them into a common linguistic form.
The distinction between a statement and the structure in which that statement acquires meaning is fundamental. If a model is asked to explain a philosophical tradition, it can often provide an accurate list of its principal concepts. Yet a philosophical system is rarely equivalent to a list of concepts. The meaning of a concept depends upon its relationship to other concepts, the questions it was intended to answer, the assumptions it rejects, the distinctions it establishes and the arguments through which it is defended. Removing a proposition from that network may leave the words intact while changing their intellectual function.
This is particularly evident when complex philosophical concepts are translated into familiar modern categories. A model may describe śūnyatā as "emptiness", ātman as "self", dharma as "duty" or karma as "action", and each translation may be useful at a basic level. Yet none of these English equivalents is identical to the conceptual field of the original term. The danger increases when the translated term is placed within a modern philosophical argument, as though the equivalence were complete. The model may consequently produce an argument that is grammatically coherent and apparently sophisticated while silently changing the conceptual conditions under which the original proposition was meaningful.
This problem is not unique to Indian philosophy. Similar difficulties arise whenever concepts travel between intellectual frameworks. "Evolution", for example, means something different when used in nineteenth-century biological theory, twentieth-century population genetics, popular social discourse or contemporary evolutionary developmental biology. "Information" has different meanings in information theory, journalism, cognitive science and ordinary conversation. "Representation" means something different in political theory, visual art, semiotics and computer science. An LLM that treats such terms as interchangeable semantic objects can produce elegant prose while obscuring the distinctions that scholarship depends upon.
The problem can therefore be described as one of relational fidelity. Factual accuracy concerns whether a statement corresponds to available evidence. Relational fidelity concerns whether the relationships among propositions, evidence and concepts have been preserved. Both are necessary.
A historical account may correctly state that an event A occurred before another event B and still be historically misleading if it presents A as the cause of B when historians have established no such causal relationship. Similarly, a philosophical explanation may accurately describe two propositions and still misrepresent a philosopher by presenting them as logically connected when the philosopher explicitly rejected that connection. An art-historical essay may identify two similar visual forms and still commit a historical error if it treats resemblance as evidence of influence.
The distinction between similarity, analogy, influence and causation is especially important for LLM-assisted scholarship. Because language models are exceptionally effective at detecting patterns of resemblance, they can easily identify conceptual parallels between otherwise distant bodies of knowledge. This capacity is potentially valuable. It allows researchers to discover questions that might not have been obvious from a conventional reading of individual sources. But resemblance is not proof of historical transmission. Two thinkers may reach similar conclusions independently, while two artistic traditions may develop similar forms in response to similar materials, experiences, or problems. Likewise, two scientific theories may use similar mathematical structures for very different reasons. A responsible scholarly argument must therefore establish what kind of relationship is actually being claimed.
Historical reasoning adds another difficulty because historians often work with incomplete or partial evidence that are sometimes contradictory. They must interpret this evidence to explain events, and they may disagree about the importance of different causes and interpretations. History is therefore not simply a matter of arranging events in chronological order; it also involves weighing evidence, assessing possible causes, and constructing explanations from what survives of the past.
A language model can produce fluent narratives, and the fluency can create an impression that historical processes were more orderly and inevitable than the evidence permits. For example, a sentence such as "industrialisation produced urbanisation, which produced new social classes, which produced modern political movements" may contain elements of historical truth. But, its linear structure can conceal feedback mechanisms, regional differences, political contingencies, and competing explanations, where the danger is not necessarily false information but excessive coherence.
This is one reason why historical scholarship requires the preservation of disagreement. If historians disagree about the significance of an event, an AI-generated synthesis Pr not automatically resolve the disagreement into a single, apparently neutral position. The disagreement itself may be part of the historical evidence, as different interpretations can reveal variations in sources, methodologies, political contexts, or conceptual assumptions. An LLM that suppresses those differences in favour of a smooth narrative may make the text easier to read while making the scholarship less accurate.
Although it does not support the stronger claim that LLMs possess one universal cultural or philosophical structure, a 2024 empirical study published in PNAS Nexus provides grounds for taking this problem seriously. In this study, PNAS Nexus compared the responses of five widely used language models with nationally representative cultural-value data and found systematic cultural patterns in their outputs. The tested models showed value patterns resembling those associated with English-speaking and Protestant European countries. The authors also found that culturally targeted prompting improved alignment for many countries. The study is important, but its findings should not be overstated: it measured cultural values using a specific survey-based framework. It tested a particular set of models, rather than demonstrating a universal "Western reasoning architecture".
A separate 2024 study on language technology identified what its authors call language-modelling bias, arguing that language technologies can represent some languages and culturally specific concepts more adequately than others. The authors connect this problem to the disproportionate emphasis on English in the development of language technologies and warn that simply adding more languages to existing systems does not necessarily capture the deeper conceptual and cultural diversity embodied in those languages. This is a more precise criticism than simply saying that LLMs are "Western": it concerns differences in representation, linguistic resources, technological development, and the conceptual assumptions built into language technologies.
A more recent research also suggests that the issue extends beyond language choice into narrative framing. A 2026 study examining GPT-4o and Claude 3.5 responses about the events of 1922 in Asia Minor found differences associated with whether prompts were presented in Greek or English, including differences in terminology, attribution of responsibility and historiographical framing. The authors explicitly acknowledge that the study uses only two models and therefore cannot establish a universal property of LLMs. Yet, it provides a useful case study for showing that the language of inquiry can affect the structure through which a contested historical event is presented.An important finding concerns the difference between factual accuracy and framing.
A 2026 pre-registered experiment with 1,912 participants found that factually accurate AI-generated historical summaries could still influence people's political views through their framing. Participants who read default GPT-4o summaries expressed slightly more liberal views than those who read Wikipedia summaries. Summaries given an explicitly liberal framing also produced more liberal responses, while conservative framing produced more conservative responses, mainly among conservative participants. The study therefore suggests that even when the facts are accurate, the way information is selected, emphasised and presented can influence how people interpret it.
This is significant development because factual accuracy alone does not guarantee neutrality. The choice of emphasis, language, context and narrative structure can shape interpretation even when the underlying information is correct.
These findings should not, however, be interpreted as evidence that LLMs are incapable of intellectual plurality. Models can produce dialectical arguments, compare competing interpretations, construct counterarguments and work with multiple perspectives when explicitly prompted to do so. The problem, instead, is that their default output often rewards compression, coherence, and completion. When a complex inquiry contains contradictions, ambiguities or several partially compatible interpretations, the model may be inclined to organise them into a more unified account because such an account is linguistically natural and useful to the user. The result is not necessarily an intentional bias; it is often the consequence of the interaction among training data, optimisation, prompting, evaluation, and the communicative expectation that an answer should be useful and complete.
This distinction matters because it changes how the problem should be addressed. If the problem were "Western bias", the solution might appear to be adding more non-Western texts to the training corpus. But intellectual diversity cannot be reduced to the number of texts or languages represented in a dataset. A tradition can be present in a corpus while its concepts are still interpreted through the categories of another tradition. Translation can preserve vocabulary while changing conceptual relationships. A text can be statistically represented while the intellectual practices surrounding that text remain absent. Diversity, therefore, concerns not only what knowledge is present, but also how relationships among forms of knowledge are represented.
This is where the idea of epistemic diversity becomes useful. A 2025 preprint studied 27 LLMs across 155 topics covering 12 countries and 200 prompt templates. For the topics tested, the models produced less diverse claims than a basic web search, although retrieval-augmented generation (RAG) increased diversity to some extent. The study also found that larger models tended to produce less diverse claims, while the effect of RAG varied by cultural context. These findings stem from a specific experimental design and should not be taken as proof that all LLMs behave the same way. They suggest that relying on a common generative system may reduce the range of information and viewpoints that users encounter.
The potential consequence is not merely that individual researchers may receive incomplete answers. If large numbers of students, researchers, writers and institutions repeatedly use the same models to formulate questions, summarise literature and generate arguments, the models become mediators of intellectual convergence. A recent review in Trends in Cognitive Sciences argues that widespread LLM use may contribute to the homogenisation of language, perspectives and reasoning by reinforcing dominant patterns and reducing exposure to alternative forms of expression and thought. This is an important emerging research direction, although it remains necessary to distinguish demonstrated effects from longer-term predictions about society.
The issue is particularly significant for artistic research because artistic practice often produces knowledge without converting it into a sequence of explicit propositions. A painting can investigate perception through colour, material and spatial organisation; an installation can investigate the relationship between body and environment through movement; a performance can produce knowledge through embodied action; and a design process can reveal possibilities through making and iteration. The knowledge produced through such practices may subsequently be articulated in language, but it is not necessarily generated by language in the first place.
An LLM working in artistic research, therefore, faces a distinctive challenge. It is extremely good at translating visual or material practices into discursive categories, but that strength can also become a form of reduction. A work concerned with ambiguity may be given a definitive interpretation. A process of experimentation may be rewritten as though it followed a predetermined methodology. A material discovery may be converted into a verbal hypothesis after the fact. The resulting academic language may be perfectly acceptable while obscuring the fact that the artwork itself constituted part of the investigation.
This problem becomes particularly acute in practice-based research, where the relationship between making, experiencing, reflecting and theorising may be iterative rather than sequential. The research question can change through practice; the method can emerge through experimentation; the outcome can modify the original question. An LLM accustomed to representing research through the familiar sequence of question, method, evidence and conclusion may reconstruct such work into a more conventional academic narrative and thereby remove the very uncertainty through which the research generated knowledge.
This does not mean that practice-based research should be insulated from analytical language. On the contrary, articulation is essential if artistic research is to enter scholarly dialogue. The important point a researcher should not overlook is that articulation should not be confused with a replacement. The written account should explain what the practice reveals without assuming that everything revealed by practice must already exist in propositional form.
The same principle applies to interdisciplinary research more generally. A productive interdisciplinary argument does not simply place several disciplines side by side. It examines the relationships among their assumptions. Neuroscience, phenomenology, visual art and artificial intelligence may all address perception. Still, they do not necessarily mean the same thing by "perception", nor do they treat the same kinds of evidence as authoritative. A superficial interdisciplinary synthesis identifies common vocabulary; a stronger one examines the differences that make comparison intellectually meaningful.
This suggests that the most useful way to understand an LLM is neither as an objective encyclopaedia nor as a defective imitation of a particular civilisation's reasoning. It is better understood as a probabilistic mediator of textual knowledge whose outputs are shaped by training distributions, model architecture, post-training processes, system instructions, prompts, retrieval mechanisms and conversational context. Its ability to produce coherent language gives it enormous value, but coherence is not the same thing as epistemic validity.
The researcher, therefore, has to distinguish at least four different questions when working with an LLM. The first is whether a claim is factually supported. The second is whether the interpretation fairly represents the relevant scholarship. The third is whether the relationships among the claims are historically, logically or conceptually justified. The fourth is whether the framing has excluded alternative positions that are necessary for understanding the subject. A response can succeed at the first level while failing at the other three.
Citations alone cannot solve this problem because an LLM may provide ten genuine references while constructing an argument that none of those sources actually support. The source can be authentic but irrelevant to the particular inference being made from it, or a quotation can be accurate but removed from its argumentative context. A secondary source can accurately report an interpretation while being mistakenly presented as primary evidence. The researcher must therefore verify not only whether the source exists but what the source actually establishes.
The appropriate response is not to reject LLMs from scholarship. Their ability to search conceptual territory, compare large bodies of literature, identify recurring themes, generate counterarguments and reveal connections can be extraordinarily valuable. What needs to change is the division of intellectual labour between researcher and model. The model should assist in mapping possibilities, while the researcher remains responsible for determining which relationships are evidentially justified.
A particularly useful practice is to ask the model to produce competing structures rather than a single synthesis. Instead of asking, "What is the relationship between A and B?", the researcher can ask: "List the possible relationships between A and B and distinguish documented influence, conceptual analogy, historical coincidence, shared intellectual context and speculative interpretation." Instead of asking for "the history of" a subject, the researcher can ask the model to separate chronology, causation, historiographical disagreement and retrospective interpretation. Instead of asking for a definitive explanation of a philosophical concept, the researcher can ask which meanings are internal to the original tradition, which are later scholarly interpretations and which are modern analogies.
Such procedures make the argument's architecture visible.
They also introduce a necessary principle of epistemic friction. A good research system should not always make knowledge easier to consume. Sometimes it should prompt the researcher to stop and ask why two propositions have been connected, what evidence supports the connection, and whether another interpretation is possible. The value of an LLM in scholarship increases when it is used not only to reduce cognitive effort but also to expose the points at which cognitive effort is necessary.
This leads to a more productive conception of AI-assisted scholarship. The objective should not be to construct a machine that eliminates ambiguity, disagreement and intellectual plurality. It should be to construct systems that can recognise those conditions and preserve them when they matter. A sophisticated model should be able to say, in effect, that two propositions are compatible but not historically connected, that a resemblance is interpretive rather than causal, that a source supports one part of an argument but not another, that a concept has no exact translation, or that scholarly disagreement cannot responsibly be reduced to a single conclusion.
Such an approach also changes the meaning of objectivity. Objectivity should not be equated with the production of one apparently neutral narrative. In many fields, greater objectivity is achieved by making competing evidence and interpretive positions visible and allowing readers to understand why they differ. A historically responsible account of a contested event is not necessarily one that eliminates disagreement; it may be one that accurately represents the grounds of disagreement.
The deeper issue, therefore, is not whether machines can think in the same way as human beings- only a philosophical speculation. The more immediate scholarly question is how computational systems transform the conditions under which arguments are assembled and circulated. If a model repeatedly takes heterogeneous material and reorganises it into similar patterns of explanation, it may gradually influence not only what people know but also how they expect knowledge to look.
That possibility deserves serious attention because intellectual diversity is not simply a matter of having different opinions. It includes differences in what counts as evidence, how propositions are related, when contradiction is productive, how context affects meaning, how uncertainty is represented and what kinds of experience are considered capable of producing knowledge. Reducing this diversity to a collection of alternative "views" would itself be a form of simplification.
The future of (responsible) LLM-assisted scholarship should therefore move beyond the narrow objectives of factual accuracy and linguistic fluency toward structural and epistemic fidelity.
A model should be evaluated not only by whether it knows the relevant facts, but by whether it preserves the relationships through which those facts become meaningful. It should distinguish evidence from interpretation, interpretation from analogy, analogy from influence, influence from causation and coherence from proof.
The most valuable LLM, in this sense, will not necessarily be the one that produces the most confident or comprehensive answer. It will be the one that can recognise when an argument requires several perspectives, when a proposition depends upon a particular context, when an apparent contradiction should be preserved rather than resolved, when evidence is insufficient and when the structure of a question itself contains assumptions that need to be examined.
The central challenge is therefore not to make artificial intelligence think in accordance with a single preferred intellectual tradition. It is to prevent the enormous diversity of human intellectual practices from being reduced, through repeated computational mediation, to a small number of highly fluent and familiar patterns.
Knowledge is not merely a collection of propositions. It is also the changing network of relationships among propositions, evidence, concepts, contexts, experiences and perspectives. If LLMs are to become serious instruments of scholarship, their development and use must increasingly recognise that architecture. The task is not simply to generate better answers, but to preserve the complexity from which meaningful questions, interpretations and arguments arise.
NB: The grammarly was used to correct spellings and grammar and edit this essay.
Further reading:
Tao, Y., Duan, N., et al. (2024). "Cultural Bias and Cultural Alignment of Large Language Models." PNAS Nexus
Shu, M., Karell, D., Okura, K., & Davidson, T. R. (2026). "How Latent and Prompting Biases in AI-Generated Historical Narratives Influence Opinions." PNAS Nexus
Papaioannou, C. (2026). "Language and the Framing of Historical Narrative in Large Language Models: The Case of Asia Minor (1922)." AI and Ethics
No comments:
Post a Comment