Direct Answer and Ethical Boundary
Using AI during a PhD literature review is not inherently dishonest. It becomes academically dishonest when a researcher conceals material assistance, submits generated text as their own unaided work, misrepresents sources that do not exist, or relies on AI output without checking whether the evidence actually supports a claim. The ethical distinction is therefore not simply human versus AI; it is transparency, source verification, intellectual ownership, and compliance with the institution’s and journal’s rules. As of 29 September 2026, many universities still lack one universal AI policy, so the researcher must examine the department handbook, supervisor agreement, ethics requirements, training guidance, and the policies of the intended journal or publisher.
Also worth reading: How Should Structural Engineers Review Responsible AI Literature for Structural Design Decisions? · How Should AI Literature Reviews Be Checked for Citation Accuracy and Research Integrity? · How Is AI Being Used Honestly in Structural Engineering Literature Reviews?
A defensible use of AI includes generating search synonyms, identifying author aliases, grouping initially relevant papers, explaining a difficult statistical method, or drafting prompts that improve database searching. A less defensible use is asking a chatbot to write a review from topic alone and treating its references as a bibliography. A prohibited or clearly unacceptable use in many settings is submitting an AI-generated review with no disclosure while pretending that every sentence and interpretation is original. The final work must still meet the university’s standard of independent scholarship, and the researcher must be able to explain the methodology, central arguments, methodological differences, and limitations of every source cited.
There is no honest percentage such as “80% AI is allowed” or “10% disclosure is sufficient.” Those numbers may appear in local guidance, but they do not transfer automatically between institutions or publishers. The most reliable rule is to obtain written agreement from the supervisor or research-integrity office when the intended use is ambiguous. Transparency is strongest when the method, tools, prompts where appropriate, verification process, and human decisions are documented.
How AI Can Help—and Where Failure Begins
AI is useful because conventional literature searching is repetitive and imperfect. A researcher may know that a paper concerns seismic response reconstruction but not whether its preferred terminology is “field reconstruction,” “response estimation,” or “structural monitoring.” A language model can propose vocabulary for a Boolean search and expose related concepts that were missing from the initial query. It can also help locate older terminology, suggest adjacent research areas, and convert a broad question into narrower combinations of database fields.
The failure begins when the model becomes an unverified authority. Generative systems can invent article titles, authors, journals, dates, DOIs, quotations, page numbers, and findings. They may combine several real studies into one nonexistent study or attach a confident conclusion to a source that says something weaker. Even when every cited work exists, an AI-generated synthesis may flatten disagreements, ignore null results, or treat correlation as causation. A fluent answer is therefore not evidence that a literature review is accurate.
Good scholarly practice treats AI output as a lead-generation system rather than a citation database. The researcher should independently open each publisher or repository record, inspect the actual paper, and record the claims that will be cited. A practical verification threshold is simple: no source enters the formal bibliography until the researcher has seen the title, authors, year, venue, and relevant primary text. Quotations require especially strict checking, including page or paragraph location. If the model is uncertain about a reference, the paper must remain outside the review unless another trusted scholarly index confirms it.
A Practical Workflow for a Defensible Review
The first step is to read the applicable policy before using AI. The researcher should identify whether the program permits AI for search assistance, coding, data extraction, text generation, translation, or manuscript preparation. It is wise to ask the supervisor in writing which uses require acknowledgment and whether prompts or outputs must be retained. A dated email or approved meeting note can demonstrate good faith, although institutional policy—not private permission—ultimately governs the researcher’s obligations.
Next, establish a reproducible search protocol. Record the databases, search date, full query, filters, date range, and inclusion or exclusion criteria. For a review begun on 29 September 2026, the search log should capture that exact date because databases are updated continuously. A broad review might begin with 150 database records, screen titles and abstracts, inspect perhaps 45 full texts, and retain 20 to 35 papers for detailed synthesis. These are illustrative counts, not universal rules; the appropriate numbers depend on the research question, review type, and available evidence.
Use AI only at labeled stages. One stage may generate ten synonym sets; another may suggest a taxonomy for coding verified articles. Researchers should never let the model choose final claims without comparison against the source. Every paragraph in the final review should be traced to at least one verified publication, and passages describing the state of research should reflect the reviewed corpus rather than the model’s general knowledge. A literature gap should emerge from documented search coverage, not simply from a chatbot’s statement that “little research has been done.”
Maintain an audit trail containing the model and version used, access date, purpose, representative prompts, outputs used, corrections made, and evidence supporting each accepted claim. Remove confidential manuscript text, peer-review material, personal data, and unpublished data before uploading it to a service. The researcher should also confirm whether prompt text is retained or used for training under the provider’s current terms. Audit records improve reproducibility, but they do not excuse a violation of privacy or intellectual-property rules.
Traditional Review Methods and AI Alternatives Compared
Traditional systematic-review methods remain the reference standard when reproducibility and bias control matter. They are slower because a person screens records, applies inclusion criteria, extracts data, and documents decisions. AI-assisted review can reduce clerical effort, but speed can create a false sense of coverage if database strategy and eligibility rules are weak. A hybrid approach generally offers the best balance: deterministic database searches and human judgment establish the evidence base, while AI assists with bounded, reviewable tasks.
| Feature | Traditional review | Unrestricted AI review | Human-verified AI-assisted review |
|---|---|---|---|
| Search reproducibility | High when fully logged | Often low | High when queries and database results are saved |
| Initial screening speed | Moderate to slow | Fast | Fast with manageable record sets |
| Risk of fabricated citations | Low | High | Low after manual verification |
| Coverage of terminology | Depends on researcher expertise | Potentially broad | Broad and iteratively refined |
| Screening bias | Low if one reviewer follows a protocol | High because the model is opaque | Lower if decisions and overrides are logged |
| Synthesis quality | Depends on domain knowledge | Fluent but unreliable | Strongest when every claim is source-checked |
| Disclosure burden | Usually none for ordinary tools | Potentially major if hidden or prohibited | Depends on institutional and publisher rules |
| Best application | Small or high-stakes reviews | Brainstorming only | Large searches with mandatory source checking |
Common Mistakes and Warning Signs
The most serious mistake is citation fabrication. A response that supplies a plausible title and DOI is not reliable merely because the DOI looks correctly formatted; identifiers should resolve to the claimed work. Another common error is “citation laundering,” which occurs when a chatbot cites a genuine paper but the supporting text was never consulted. Reviewers may detect this through mismatched terminology, impossible quotations, or claims that do not match the paper’s method.
Researchers also err by reviewing only studies surfaced by a conversational model. Chatbots do not consistently expose database indexing, review methodology, retraction status, or complete result counts. The model may privilege recent popular sources, English-language publications, or well-indexed journals. Those biases matter in structural engineering, where codes, local seismic conditions, construction practices, material standards, and building typologies vary by region.
A second warning sign is an unexplained change in writing style. A polished section that contains too many generic transitions but no specific methods, authors, datasets, or limitations may be machine-generated. A more practical issue is untraceable interpretation: if the researcher cannot say why one study supports a category while another does not, the synthesis is not ready for submission. Quantity is also misleading; reviewing 300 titles but reading only abstracts is not equivalent to examining 30 full papers.
Finally, confidentiality and authorship must be managed deliberately. AI should not be listed as an author merely because it drafted or edited text, because authorship requires accountability and contribution criteria. Using AI to process identifiable interview data or confidential industry drawings can create governance problems even if no manuscript is submitted. The safest process separates public source material from restricted data and applies institutional controls before external processing.
Costs, Tools, and Practical Thresholds
The monetary cost ranges from zero to substantial, depending on the service and the scale of the work. A researcher can begin with free database access, institutional library licenses, a general-purpose chatbot’s accessible tier, and manual reference management. Paid subscriptions commonly cost from roughly US$20 per month for an individual general AI plan to several hundred dollars per month for business or enterprise access, but prices change by provider, region, and usage limits. University library access may include databases, statistical tools, or licensed AI products at no additional personal charge.
The more important cost is time. A 20-paper review with individual extraction, cross-checking, and synthesis may require 80 to 150 hours after the search is complete, although a complex systematic review can take much longer. AI can shorten clerical stages by perhaps 20% to 50%, but those estimates are workflow observations rather than guarantees. It may not reduce total time if prompts must be debugged, references manually verified, or contradictory outputs investigated. Institutions should budget for software licenses, training, secure storage, and supervision rather than assuming that generative AI makes rigorous review free.
Choose a tool according to task and risk. General chatbots are useful for conceptual orientation, but bibliographic databases and publisher platforms remain better for reproducible discovery. Reference managers organize records; they do not independently prove that a paper supports a claim. Translation models can assist with non-English sources, yet the original text should be checked where an exact interpretation matters. Before adoption, test a tool on 5 to 10 known papers: see whether it identifies their core claims, misses the same concepts consistently, fabricates references, or can distinguish study quality.
When to Act and What to Disclose
Disclosure should be written specifically enough for a reader to understand the contribution. A useful statement identifies the tool or tool class, the tasks performed, the date range of use, and that outputs were manually checked. It should not claim “AI was used only for grammar” if it was also used to propose search terms or classify articles. The disclosure should remain compatible with privacy requirements, especially if a prompt contains unpublished material or data governed by an agreement.
The researcher should pause and seek formal guidance when the intended use includes generating substantial passages, interpreting confidential peer-review comments, extracting sensitive data, translating evidence used in a safety decision, or preparing text for a journal with a strict policy. It is also appropriate to ask when the evidence base may influence structural design, public safety, or policy. Current work cited for this site includes research on AI-assisted structural realignment and AI-driven reconstruction of structural responses; those applications are technically promising, but they do not transfer a literature-review use case into full design authority.
A prudent decision rule is to use AI when the task is reproducible, inspectable, and low risk; prohibit its use as a substitute for scholarly judgment. Ask, “Could another researcher repeat this process and obtain the same evidence?” If the answer is no, improve the process. Ask, “Can I point to the primary source behind every consequential statement?” If not, remove or verify it. Ask, “Will I disclose this assistance in accordance with local rules?” If not, obtain approval before proceeding.
The Defensible Standard for AI Structural Engineering Research
The strongest position in September 2026 is neither blanket acceptance nor blanket rejection. AI can reduce repetitive searching and help engineers navigate unfamiliar terminology, especially across structural health monitoring, seismic assessment, computational civil engineering, cost prediction, and resilience. It can also make a review less repetitive. But it cannot be trusted to establish what a body of structural engineering literature collectively proves, identify all relevant local evidence, or replace transparent scholarly reasoning.
For a PhD literature review, a defensible submission should contain a logged search strategy, transparent eligibility criteria, verified primary sources, a clear synthesis of agreements and disagreements, and a disclosure consistent with institutional and publisher rules. The researcher must retain intellectual ownership: not only correcting the model, but explaining why sources are relevant and how their evidence changes the argument. If AI was used materially, stating that plainly is preferable to hiding it and losing credibility later.
Thus, “Is using AI for a PhD literature review dishonest?” has a conditional answer: no when it is a disclosed, bounded aid under the researcher’s control; yes in many academic-integrity terms when it substitutes for undisclosed reading, fabricates evidence, conceals substantial authorship assistance, or violates an explicit rule. The decisive test is whether the review remains reproducible, source-grounded, transparent, and recognizably the scholar’s own work.