CustomGPT.ai Blog

Using AI for Grant Writing: What Funders Require

Author Image

Written by: Alden Do Rosario

·

31 min read

AI drafting is allowed; unverified claims are what funders sanction

No major U.S. funder bans AI drafting. NIH’s NOT-OD-25-132 states that “NIH will not consider applications that are either substantially developed by AI, or contain sections substantially developed by AI, to be original ideas of applicants” — an originality standard, not a disclosure form.

NSF’s December 2023 Notice to the Research Community says proposers “are encouraged to indicate in the project description the extent to which, if any, generative AI technology was used.”

NSF 26-200, Supplement 1 to PAPPG 24-1, makes fabrication, falsification, or plagiarism actionable as research misconduct whether committed directly or “through the use or assistance of other persons, entities, or tools, including artificial intelligence (AI)-based tools.”

DOE’s standard Notice of Funding Opportunity Part 2, Section II.B.3, requires that any AI use in creating any part of an application be appropriately attributed.

Three agencies, four instruments, one common target. The funder is not asking whether you used AI. The funder is asking whether you can prove a human verified it.

NOT-OD-25-132 addresses AI-drafted application content rather than factual accuracy, and it states that if the detection of AI is identified post award, NIH may refer the matter to the Office of Research Integrity while “taking enforcement actions including but not limited to disallowing costs, withholding future awards, wholly or in part suspending the grant, and possible termination.”

The working definition transfers cleanly from privacy review: a verified response is one you can defend later — right audience, right data, right purpose, and an auditable evidence trail behind it. In a grant narrative, every factual claim traces to a named source document and a qualified human signed off on the trace.

The exact readings, agency by agency, sit under “What each funder actually requires, notice by notice.”

The exposure is fabrication, not detection

Most published advice on AI grant writing asks whether reviewers can tell. That is the wrong exposure. No examined funder policy names detector output as a sanction trigger.

Detectors are unreliable and reviewers do notice flattened voice, but NIH’s originality rule turns on whether the ideas are the applicant’s, not on whether the sentences read as machine-written, and NSF 26-200 turns on fabrication, falsification, and plagiarism. Neither hinges on stylistic detection. No amount of prose-polishing addresses either.

The documented failure mode is worse than the folk version of it. A language model asked for supporting literature assembles citations whose parts are individually real for works that were never published.

Walters and Wilder, measuring bibliographic citations across 42 topics in Scientific Reports, found “55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated” — and that “most of the fabricated article, book, and website citations include the names of real journals, publishers, and organizations,” with only 5% of fabricated works also inventing the larger work they appeared in.

Read that the right way round. Confirming that the journal exists, or that the authors are real people, fails to catch nearly every fabrication. The fabricated citation is assembled out of real components, so the components all check out. The only thing that does not exist is the specific work — the one thing a spot-check of the parts never tests.

Under NSF 26-200 that is fabrication committed through an AI-based tool, and the tool is not a defense. The countermeasure is architectural rather than editorial: grounding responses in your own documents and enforcing citation discipline removes the conditions under which a citation gets invented in the first place.

“Verify every claim” does not survive a deadline

“Verify every claim” is the standard advice and it is correct. It is also unimplementable as stated. On a Specific Aims section drawing on sixty sources, under a deadline, a human instruction to check everything produces a checked sample and an unchecked remainder — with no record distinguishing the two.

The imperative exists everywhere; the mechanism exists nowhere. Published guidance issues “verify all data” and stops: no workflow, no cost, no reviewer handoff, no artifact at the end. When the question arrives eighteen months later at a post-award audit, “we checked” is a memory, not a record.

So the design question is not how to draft faster. It is how claims arrive already linked to their source, with the unlinked ones marked. Inverting the workload — verify by exception rather than verify everything — is the only version that survives a deadline, and compliance reviewers reached that conclusion long before grant offices did.

Their durable form of proof is linking every factual claim to the exact document, section, or snippet it came from, with a version reference attached. That inversion is what claim-level verification does.

Diagram showing an AI response split into four separate factual claims: two connect by solid lines to specific source documents and are marked verified, while two connect to nothing and are flagged as untraced, producing a Verified Claims Score of 50 percent — a measure of traceability to your sources, not of truth.

CustomGPT.ai’s Verify Responses extracts every claim and traces it to a source document

Verify Responses breaks an AI response into its individual factual claims and cross-references each one against your uploaded source documents, showing the filename, page number, or URL where each claim was found. Claims with no traceable support are flagged. The result is a per-response record of what is evidenced and what is not.

Chat interface showing CustomGPT claims being verified with a 92% accuracy label and highlighted text.

Four components handle extraction, scoring, stakeholder review, and analytics

The Claim Verifier “automatically extracts every factual claim from a response and cross-references it against your source documents.” The Verified Claims Score reduces that to a number. The Trust Score runs the response past a virtual committee of six stakeholders. Customer Intelligence Integration pushes the results into the analytics dashboard, where conversations can be filtered by score.

Those are the conceptual names. The labels you click are different: the verification pop-up presents a Claims tab showing extracted claims with source verification and the verified claims score, and a Trust Building tab showing stakeholder satisfaction status and recommendations. What makes the output usable in a research office is that verification surfaces the actual source location for each claim — filename, page number, URL — rather than a global confidence number with nothing underneath it.

The Verified Claims Score is verified claims divided by total claims — and it measures traceability, not truth

The formula is arithmetic, not a model output dressed up as a metric. It is the Verified Claims Score, which divides verified claims by total claims found in the response. Ten factual statements with eight traceable to your documents scores 80%.

Now the part that matters more than the formula. A high Verified Claims Score means the claims in the response can be traced back to your source documents. It does not guarantee the source documents themselves are correct or complete. The feature verifies alignment with your sources, not absolute truth.

For a grant office the stake is specific: if the stored biosketch, budget justification, or preliminary-data summary is stale or wrong, a 95% score certifies faithful reproduction of a wrong document. None of that removes the principal investigator’s signature from the application, which is why keeping a named human accountable for the claims that go out the door remains the load-bearing control.

CustomGPT.ai’s own guidance on thresholds is risk-tolerance advice, not product behavior: for high-stakes applications such as legal advice, medical information, or compliance, scores above 90% are suggested, while general customer support may accept 80% or better.

Nothing in the product blocks a submission at any score. The threshold is the office’s choice and the office’s responsibility.

Verify Responses Claims dashboard showing verified claims, source verification metrics, claim statuses, and an AI-generated accuracy disclaimer.

Six stakeholder perspectives read the same response for different risks

The Trust Score runs a response through a simulated review by six virtual stakeholders, each asking a distinct question. End User: is this helpful and appropriate? Security/IT: any data exposure or system security concerns? Risk Compliance: any regulatory or financial exposure?

Legal Compliance: any liability risk or legally ambiguous language? Public Relations: could this cause negative public perception? Executive Leadership: does this align with brand and strategic goals?

Two of those map directly onto grant reality. Legal Compliance is the lens that catches a phrase overstating what preliminary data supports. Executive Leadership is roughly the lens an institutional signing official applies before their name goes on the submission.

The two scores can disagree, and that disagreement is by design: a response can be factually traceable and still raise a stakeholder concern, because a statement might be true but use legally risky phrasing. The system highlights the conflict rather than resolving it. Resolution is a human’s job.

Trust Building tab showing an overall Approved status across six stakeholder perspectives: End User, Security IT, Risk Compliance, Legal Compliance, Public Relations, and Executive Leadership.
Example table showing four extracted claims, two verified against the source and two flagged for review.

Take a Specific Aims sentence: “Our pilot cohort of 42 participants showed a 31% reduction in readmission at 90 days, consistent with Chen et al. (2024).” That sentence contains four separate factual claims. CustomGPT.ai’s Verify Responses splits them, checks each against the uploaded corpus, and returns a different verdict for each one.

Two of the four claims trace to a source and two do not

This is a worked illustration built on a synthetic corpus. The documents, the numbers, and the missing reference are invented to show the mechanic, not lifted from a real submission.

An agent’s knowledge base holds a pilot study report (pilot-readmission-study-v3.pdf), the IRB protocol, two prior progress reports, and the lab’s published papers. A drafter asks for a Specific Aims paragraph, the response contains that readmission sentence, and the builder clicks the shield icon next to that response and selects “Run for this response.” Results appear in seconds.

Extracted claimStatusSource locationReviewer action
Pilot cohort of 42 participantsVerifiedpilot-readmission-study-v3.pdf, p. 4Accept, no action.
31% reduction in readmissionVerifiedpilot-readmission-study-v3.pdf, p. 11, Table 3Accept, but confirm Table 3 reports the same endpoint the sentence claims.
At 90 daysUnverifiedNo source locatedTable 3 covers 30-day readmission. The 90-day figure is not in the corpus. Locate the 90-day analysis and add it to the knowledge base, or correct the sentence to 30 days.
Consistent with Chen et al. (2024)UnverifiedNo source locatedNo Chen 2024 exists in the uploaded papers. Resolve the reference against the actual literature before it goes near a submission.

An unverified claim is untraced, not false, and only a human can tell which

Two of four claims traced, so the Verified Claims Score is 50%. The flags are not findings of falsehood, and reading them that way produces the wrong next action. An unverified claim could not be traced back to the uploaded documents; that is not the same as being wrong.

Claim 3 may well be true and simply live in a file nobody uploaded — a postdoc’s 90-day analysis in a folder outside the corpus. Claim 4 may be a real paper the corpus does not contain, or a fabrication with a DOI that resolves to nothing. The system does not distinguish those two cases. It cannot. The human does, and that distinction is the entire reason a named reviewer is still the gate rather than a formality bolted onto a score.

The Verified Claims Score here is 50%, and that is the right answer

Fifty percent is not a failing grade. It is an accurate description of a draft sentence at an early stage, and the useful part is not the number but the routing. The reviewer’s attention went to two claims instead of four, and the two that got attention were the two that needed it.

What exists afterward is the second half of the payoff: a record of which claims were reviewed, against which sources, at which page numbers. Verify Responses provides audit trails but requires human judgment for root-cause fixes, so the record is evidence rather than a remedy.

Claim 4 is precisely the shape of thing NSF 26-200 treats as fabrication if it reaches a submitted proposal unchecked, and “the AI produced it” is not a defense.

CustomGPT.ai’s Verify Responses runs on one response or on all of them, and it costs credits either way

A builder either turns verification on globally so every response is analyzed automatically, or clicks the shield icon beside a single response — including one in a past conversation — to run it once. Verification happens after the response is generated, so end users see no added delay, and they never see the panel at all.

CustomGPT.ai Ask interface highlighting the shield icon beneath an AI response used to run Verify Responses.
A 78-second walkthrough of turning on and reading CustomGPT.ai’s Verify Responses.

Global verification analyzes every response; the shield icon runs one on demand

The global path, per the documentation, is: click the three dots menu on the dashboard, click Actions, toggle Verify Responses On, then click the “I understand” button to confirm. The per-response path skips global enabling entirely — click the shield icon next to any AI response and select “Run for this response” when prompted.

Retroactive runs work on any past conversation, with results in seconds, and that matters more in grant work than in support work. A research office that wants to audit a narrative drafted six weeks ago does not need to have decided in advance to verify it.

Verification costs query credits and only the builder sees the results

Verification uses additional query credits when enabled. Any advice to simply leave it on globally without naming that cost is incomplete.

The cost shapes the sensible operating pattern rather than arguing against it. Run verification on everything while drafting, when volume is low and the value of finding gaps is highest. Move to per-response spot checks plus periodic audits once the narrative is stable and the evidence base has stopped changing.

Visibility is narrower than most teams expect. Verify Responses is a behind-the-scenes tool for the builder and administrator; end users see a standard chat interface with no shield icon and no panel.

Analysis is performed securely within the CustomGPT.ai environment, verification data is not shared with other users or external systems, and only the authenticated user can access the analysis for their queries.

A low score usually indicts the evidence base, not the AI

When CustomGPT.ai enabled Verify Responses on its own support agents and tracked every response, the finding was not that the AI was failing:

The AI wasn’t always the problem. Our documentation was. We discovered gaps we never knew existed – questions users were asking that our knowledge base didn’t fully cover. The AI was doing its best with incomplete information. And “doing its best” meant guessing. — Introducing Verify Responses, the CustomGPT.ai launch post

The remedy was equally unglamorous: fixed persona configurations, resolved system-level retrieval issues, and two new documentation articles that directly addressed user needs.

That maps onto grant work without strain. A Specific Aims section that scores low is usually not evidence of a fabricating model. It is evidence that the claims the narrative wants to make outrun the documents the office has assembled. Same diagnosis, larger consequence.

Every inaccurate response is a signal. It’s telling you something. Either your AI needs tuning, your system needs fixing, or your content has gaps. You just have to listen. — Marko Mitrović, Product Manager at CustomGPT.ai

The three-way triage is the transferable part. A low score sorts into persona configuration, the retrieval system, or gaps in documentation, and in a research office the third bucket dominates. The 90-day analysis lives in a postdoc’s folder.

The updated biosketch was never uploaded. The progress report cited in the narrative is two versions behind. Low scores point directly at content you need to create, which makes the feature a knowledge-gap finder at least as much as a checker.

What each funder actually requires, notice by notice

Three agencies, four instruments, one common target. NIH’s NOT-OD-25-132 tests originality. NSF’s December 2023 notice encourages disclosure, and NSF 26-200 makes fabrication through an AI tool research misconduct. DOE’s standard NOFO Part 2 mandates attribution.

Foundations and state agencies mostly publish nothing and enforce the same rule through the accuracy representation on the application form. Each reading is exact, because the differences between them are where teams get the rule wrong.

NIH NOT-OD-25-132 is an originality rule, not a disclosure regime

The operative sentence in NOT-OD-25-132 is narrow and worth reading exactly: “NIH will not consider applications that are either substantially developed by AI, or contain sections substantially developed by AI, to be original ideas of applicants.” That is a test of whose ideas the application contains.

The notice adds no AI checkbox and no disclosure form, and pages that characterize it as “disclosure, not prohibition” have it backwards.

The notice carries a single effective date covering everything in it: “This policy is effective for applications submitted to the September 25, 2025, receipt date and beyond.”

Both of its rules sit under that date — the originality standard, and a companion cap under which NIH “will only accept six new, renewal, resubmission, or revision applications from an individual Principal Investigator/Program Director or Multiple Principal Investigator for all council rounds in a calendar year.”

NIH tied the cap partly to rising volumes of AI-generated applications. Check the exception list against the current NIH Grants Policy Statement §2.3.7.12 rather than the originating notice — as revised March 2026, the policy “applies to all activity codes with the exception of T (Training grants), R25 (Research Education Grants) and R13 (Conference Grants),” a wider carve-out than the notice itself named.

No form is not the same as no expectation. NIH’s May 14, 2026 Extramural Nexus reminders on AI and research integrity ask researchers to “clearly describe in applications, manuscripts, and presentations the use of the AI tools” and how they were used — narrative, not a form field. It sits alongside NIH’s statement that NIH and ORI “work together” when possible misconduct arising from these tools surfaces, and it restates the originality standard and the cap without changing either.

NSF’s disclosure standard is voluntary, and its misconduct rule is not

NSF’s standard reads in full: “Proposers are encouraged to indicate in the project description the extent to which, if any, generative AI technology was used and how it was used to develop their proposal.” Encouraged, not required.

The notice attaches no stated penalty to a proposer who declines — but that is silence rather than a safe harbor, and the same notice holds proposers “responsible for the accuracy and authenticity of their proposal submission in consideration for merit review, including content developed with the assistance of generative AI tools.” Low disclosure stakes, unchanged accuracy stakes.

The disclosure ask is not in the Proposal and Award Policies and Procedures Guide’s proposal-preparation instructions — it is a standalone December 2023 notice. NSF’s published version list shows the current guide is PAPPG 24-1, applying “to all proposals submitted or due on or after May 20, 2024,” now modified by two supplements that take precedence over it: NSF 26-200, effective December 8, 2025, and NSF 26-202, effective January 22, 2026. That list runs 24-1 straight to a deferred 26-1.

NSF’s own explanation is that “the release of Executive Order 14332: Improving Oversight of Federal Grantmaking requires the Office of Management and Budget (OMB) to streamline and transform the Uniform Guidance,” and so “to ensure alignment with the Uniform Guidance, NSF will defer release of NSF 26-1 PAPPG.” Anyone citing a “PAPPG 25-1” is citing a document NSF never published.

The rule that changed for 2026 is NSF 26-200, Supplement 1 to PAPPG 24-1, effective December 8, 2025. It expands the Chapter XII.C definition: “RESEARCH MISCONDUCT means fabrication, falsification, or plagiarism, whether committed by an individual directly or through the use or assistance of other persons, entities, or tools, including artificial intelligence (AI)-based tools.” Using a tool is not a defense.

The conduct attaches to the person. Voluntary disclosure changes nothing about the verification burden.

NSF’s confidentiality ban binds reviewers, but the accuracy duty is the applicant’s

NSF’s notice also states that “NSF reviewers are prohibited from uploading any content from proposals, review information and related records to non-approved generative AI tools,” and that doing so counts as a violation of the agency’s confidentiality pledge and other applicable laws, regulations and policies.

As written, that prohibition binds reviewers, and NSF has not extended it to proposers. It has not left the applicant side unaddressed either — the same notice encourages proposers to describe their generative AI use and holds them responsible for the accuracy and authenticity of the submission, AI-assisted content included. That accuracy duty is stated policy and it applies to you.

Applicant-side confidentiality is a separate matter, and an inference rather than an agency position: the disclosure event is structurally identical when unpublished specific aims go into a public chatbot. The proposal content leaves the institution’s control either way, and a system that keeps the corpus inside a controlled environment has a specific answer to that.

CustomGPT.ai is listed as GDPR compliant, states that customer data is not used for model training, and holds SOC 2 Type 2 certified security controls — controls, not compliance guarantees. No deployment is automatically compliant because a vendor holds certifications.

Teams inside agency procurement and review constraints usually start from the public-sector deployment path federal agencies typically follow rather than a consumer chatbot account.

DOE’s NOFO template carries the only genuinely mandatory attribution duty

The DOE Financial Assistance Notice of Funding Opportunity Part 2 language at Section II.B.3 is unambiguous: “Any use of artificial intelligence in the creation of any part of an application for this NOFO must be appropriately attributed.” That is stronger than NSF’s voluntary ask, and across NIH, NSF, and DOE it is the only mandatory AI attribution duty.

It is NOFO-template language rather than department-wide policy — the text says “for this NOFO,” and it lives in the Part 2 companion document, which DOE describes as containing “fixed DOE requirements that generally do not change from NOFO to NOFO.” It recurs as boilerplate, but its force is per-opportunity, so verify the wording and section numbering against the Part 2 attached to your specific opportunity rather than the copy cited here.

And DOE P 203.1, “Use of Generative Artificial Intelligence,” is not the document to cite despite its authoritative-looking title and recency: its own scope states that “the policy, as implemented by DOE programs, does not apply to recipients of financial assistance from DOE” and “does not apply to the use of GenAI to carry out basic research or applied research.”

Foundation and state funders enforce the same rule through the representation you sign

DOE’s operative mechanism is representational rather than punitive: “Even with the use of artificial intelligence, each applicant is responsible for and is representing to the U.S. Government that the information in its application documents is accurate, that the applicant is fully capable of performing the work described in the application, and that the submission of the application does not and will not infringe or violate any rights of any third party or entity.”

No AI-specific sanction appears in that section, which states a duty rather than a penalty. Inaccurate AI-generated content draws the general exposure attached to any misrepresentation — the point being that it needs no AI-specific rule to bite.

A missing attribution is a separate failure with its own ladder: the Part 2 section titled “Requirement for Full and Complete Disclosure” reaches cancellation of award negotiations, modification or suspension of a funding agreement, debarment proceedings, and civil or criminal penalties for failure to disclose requested information. Accurate output does not cure an unattributed draft.

Most private foundations and state agencies publish no AI policy at all. The absence is not permission. It is silence around an unchanged rule, because every application form asks the applicant to affirm accuracy.

A grant writer who cannot show where the outcome number, the beneficiary count, or the prior-award figure came from has the same problem as a principal investigator citing a study that does not exist, minus the notice number.

The standard generalizes even where the sanction does not. The operative rule for anything unsupported is to treat the claim as unverified and fall back to safe responses. No source, no claim.

A nonprofit’s evidence base is annual reports, the 990, and the logic model

A small nonprofit’s corpus differs in content and is identical in function: the last three annual reports, the program evaluation, the 990, the logic model, prior successful applications. Those are the documents a grant narrative’s claims should trace to, and they are the same kind of artifact as a lab’s preliminary data.

Budget is a real constraint on that answer, and a two-person development shop has one where a research university does not. What a grant-funded team can realistically afford to run narrows the field before any feature comparison starts, and pretending otherwise helps nobody.

Research administrators need an institutional rule they can defend, not a per-PI habit

The person routing forty proposals through a submission deadline has a different problem from the principal investigator drafting one. You need a rule that applies across labs, produces the same evidence every time, and holds up when a program officer asks eighteen months later how a specific number got into a specific narrative.

An institutional AI rule specifies scope, corpus, attribution, and verification

An institutional rule that survives contact with reality specifies four things, and each maps to a notice rather than to a preference. Scope: which sections AI may assist with, given NIH’s originality standard, which turns on whether the ideas are the applicant’s.

Corpus: drafting happens against approved institutional sources, never a public chatbot with unpublished aims pasted in, given NSF’s confidentiality framing. Attribution: a standing practice of naming the tool and its use in the project description, which satisfies NSF’s voluntary ask by default and DOE’s mandatory attribution when a DOE Part 2 applies. Verification: a named human reviewer per submission, and a record of what they reviewed.

That is a defensible structure, not legal advice. Your own counsel and signing officials own the final rule, and they should — the four elements are a starting frame, and the institution carries the consequence.

The evidence packet is per-claim source citations plus a stakeholder flag

CustomGPT.ai’s Verify Responses creates audit-ready documentation for every AI response: the Legal Compliance stakeholder flags liability risks, and verified claim scores supply source citations. Where “the AI said so” is not a valid answer, that is the proof available.

Verification data flows into the Customer Intelligence dashboard, where conversations can be filtered by verified claim score or stakeholder status. For an administrator, that filter is the queue — narratives scoring below the office’s own threshold, surfaced before a deadline rather than after an award.

An application is questionnaire-shaped, which is why the closest existing analogue is a compliance questionnaire: map each question to the controls and evidence that answer it, then generate answers that arrive with citations attached.

Alternatives exist, and the useful question is what evidence each one leaves behind

Purpose-built grant tools are worth evaluating. Grantable, OpenGrants, Fundrobin, and Thesify all work this space. The distinguishing question is not who drafts better prose. It is what evidence each one leaves behind.

Thesify is genuinely strong on policy. Its treatment of NIH’s AI rules works through NOT-OD-25-132 and adjacent notices in detail and publishes a documentation table covering tool name and version, date of use, proposal section, prompt boundary, human reviewer, verification step, and final decision. That table is good practice and worth adopting regardless of tooling.

The gap is that every remedy it prescribes is manual: its stated validation loop is that “every author string, journal title, and DOI must be manually cross-checked against independent indices like PubMed prior to routing.” That documents the check without changing the odds the check is needed, and it never asks whether the drafting system should make fabrication structurally harder in the first place.

Grantable reads the notices differently. Its coverage of the “can funders tell” question states that “NIH requires that applicants disclose the use of AI tools in their applications,” and that “NSF’s policy focuses on transparency and proper attribution.” The NIH half does not survive contact with the notice.

NOT-OD-25-132 states an originality standard and asks for no disclosure at all — a difference that matters, because a team optimizing for a disclosure it does not owe will skip the verification it does.

Accuracy language is worth reading precisely, in both directions. The zero-hallucination pitch is rarer among these tools than the category’s reputation suggests:

Fundrobin describes grounded AI with “low hallucination rates”, a measured and falsifiable claim, and OpenGrants’ 2026 playbook is blunter than most vendor copy, warning that “AI fabricates sources at a rate that has not meaningfully improved.” Where an accuracy guarantee does appear with no error budget attached, a research office should treat it as a reason for more diligence, not less.

The same standard applies to CustomGPT.ai. Its controls — a context boundary that walls the AI to supplied data, retrieval-augmented generation, and citation support — reduce hallucination sharply without eliminating it, and promise no perfect accuracy in every edge case.

That is a weaker claim than zero hallucinations. It is also the one that is true. The useful comparison across all of these is whether output arrives with a per-claim source location and a reviewable record, or whether verification remains a human instruction bolted on afterward.

What CustomGPT.ai’s Verify Responses does not do

Verification is a control, not a guarantee. Three limits set the boundary:

  • It is builder-only by design. End users never see the panel, which makes it a pre-submission control, not a live disclosure to a reviewer.
  • It runs after generation rather than blocking it. Nothing is withheld while the check runs, and nothing is stopped by a low score.
  • The scores are AI-generated and work best as a guide, not as a definitive measure.

Starting with one narrative and one corpus

The smallest useful test is one agent, one grounded corpus, and one draft section. Pick a single Specific Aims section or one foundation letter of inquiry. Upload only the documents its claims should trace to. Generate it. Click the shield icon. Read the unverified list.

That list is the real deliverable on day one. It is a to-do list for the evidence base, not a report card on the AI, and the gap it exposes between what the narrative asserts and what the office can evidence was there before any model got involved.

The terms are worth knowing before you start. CustomGPT.ai’s Verify Responses runs on Standard, Premium, and Enterprise — verification consumes additional query credits on any plan, and burns fewer of them on Premium and Enterprise.

The free trial runs 7 days and is cancellable anytime, a credit card is required to sign up, and after the 7-day trial ends the card is automatically charged for the plan selected at signup. Teams with procurement, security review, or volume requirements take the enterprise path instead.

The funder is not asking whether you used AI. It is asking whether you can prove a human verified it. That question has an answer only if something recorded the check.

FAQ

Is it against NIH policy to use AI to write a grant application?

No. NIH does not ban AI drafting. NOT-OD-25-132 states that “NIH will not consider applications that are either substantially developed by AI, or contain sections substantially developed by AI, to be original ideas of applicants.” That is an originality standard, not a prohibition and not a disclosure form. The notice carries one effective date covering everything in it — the originality standard and the six-application cap alike — for “applications submitted to the September 25, 2025, receipt date and beyond.” Using AI to draft is permitted; submitting an application whose ideas are the model’s is what fails.

What does “substantially developed by AI” actually mean?

NIH has not defined it. NOT-OD-25-132 supplies the phrase and no percentage, word count, or section test. Anyone quoting a threshold is inventing one. The defensible posture is not to guess where the line sits but to make the scientific argument demonstrably human-authored — your hypothesis, your interpretation, your preliminary data — and to keep a record showing which sections a human wrote and verified.

Do I have to disclose that I used AI on my grant application?

It depends on the funder, and the three federal agencies differ. NIH requires no disclosure — NOT-OD-25-132 tests originality instead. NSF’s December 2023 notice says proposers “are encouraged to indicate” AI use in the project description: voluntary, with no stated consequence for declining. DOE’s Financial Assistance Notice of Funding Opportunity Part 2 is the strict one — attribution is mandatory for that opportunity.

Can grant reviewers tell if you used AI?

Sometimes, and it is the wrong thing to worry about. AI-detection tools are unreliable, and no examined funder policy names detector output as a sanction trigger. Program officers reading dozens of proposals are not running detection passes. A flattened voice may cost you on impact, not integrity. The exposure that carries an actual penalty is different: a fabricated citation is fabrication, and NSF 26-200 makes that research misconduct.

What happens if my grant proposal cites a study that doesn’t exist?

That is fabrication, not a formatting error. NSF 26-200, Supplement 1 to PAPPG 24-1, defines research misconduct as “fabrication, falsification, or plagiarism, whether committed by an individual directly or through the use or assistance of other persons, entities, or tools, including artificial intelligence (AI)-based tools.” The tool is not a defense — the conduct attaches to the person who signed. A plausible author string with a DOI that resolves to nothing is the classic failure mode.

Can I get in trouble for using AI after the grant is already awarded?

Yes. The risk does not expire at submission. NOT-OD-25-132 states that NIH may refer a matter to the HHS Office of Research Integrity while taking enforcement actions “including but not limited to disallowing costs, withholding future awards, wholly or in part suspending the grant, and possible termination.” A claim nobody verified in month one is still unverified when a program officer asks about it in month eighteen.

Is it safe to paste unpublished specific aims into a public chatbot?

Treat it as a disclosure event. NSF’s notice prohibits reviewers from uploading proposal content to non-approved generative AI tools, calling it a confidentiality violation. That rule binds reviewers — NSF has not extended the confidentiality prohibition to applicants, though it does address the applicant side elsewhere in the same notice, encouraging proposers to describe their AI use and holding them responsible for the accuracy and authenticity of the submission. The content leaves your institution’s control either way, which is why research offices typically draft against a controlled environment rather than a consumer account.

How do I word an AI disclosure statement for an NSF proposal?

NSF mandates no boilerplate. Its December 2023 notice sets an extent-and-manner standard, so a statement that satisfies it names three things: the tool and version, what it did and to which sections, and confirmation that a qualified human verified the output. For example: “Generative AI [tool, version] was used to draft language in Section X from the investigators’ own source documents; all factual claims and references were verified by [name].”

What should I document about AI use before I submit?

Enough to reconstruct the check later. A workable record names the tool and version, the date, the section, what the tool was asked to do, what data went in, who reviewed the output, and what they verified it against. Documentation does not prove compliance by itself — the useful question is which of those fields your drafting system logs automatically versus which a person retypes into a spreadsheet from memory.

Can AI find real studies and statistics for me to cite?

Not by recall. A language model asked to remember literature predicts what a citation looks like, which is exactly how phantom papers get plausible authors and dead DOIs. Retrieval over a real corpus is a different operation: searching PubMed, NIH RePORTER, or your own uploaded papers returns records that exist because a database returned them. Same-looking request, opposite risk profile. Use retrieval to find sources and never a model’s memory.

What if an AI detector falsely flags my grant application?

You cannot argue with a detector, and false positives fall hardest on researchers writing in a second language: Liang et al., Patterns (2023) found that GPT detectors “consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified.” What you can produce is evidence: drafts showing your authorship, and a record of which factual claims trace to which source document, at which page. A provenance trail answers the question a detector only raises. Build it while drafting — reconstructing one after a flag arrives convinces nobody.

Does a high Verified Claims Score mean the proposal is safe to submit?

No, and reading it that way is the most consequential misread. A high Verified Claims Score — the ratio CustomGPT.ai’s Verify Responses reports — means claims trace back to your source documents; it does not guarantee those documents are correct or complete. If a stored biosketch or preliminary-data summary is stale, a 95% score certifies faithful reproduction of a wrong document. The scores are AI-generated guidance, no product gate blocks a submission, and the principal investigator still signs.

Build AI agents from your content, in minutes!