FAQ

How the platform uses AI

What is sent to AI services, what is not, how retention is prevented, and how AI-reconstructed figures are flagged.

Status: draft, under review. The canonical version of this text is published on the public site at nia-lp-marketing.vercel.app/legal/ai-processing; this page mirrors it for readers inside the documentation.

This page explains, in plain terms, where the platform uses AI, what leaves the platform when it does, what never does, and what stays your responsibility. It describes the platform as built; where a setting depends on how Nia has configured its environment we say so.

The short version

  • AI on this platform proposes; people decide. Extracted indicators, reconstructed figures and suggested actions are all reviewed by a person before they count. Advisor answers are shown to you straight away, with their citations, as advice for you to check — they never change your data.
  • Spreadsheets, CSV, Policy Hub documents and Advisor attachments are always read on the platform. PDF, image and scanned files, Word files with no readable text, and any PDF or Word data submission in which the platform's own reader finds no reportable figures are sent to the external parsing service (section 2). Only if the platform's own reader finds nothing in a sheet may a bounded portion of that sheet's text be sent to an AI model (section 1).
  • PDFs, images and scanned files are sent to one external parsing service, with caching switched off.
  • Text sent to AI models goes only to named hosts that do not store or train on it; for conversational and extraction models, only to hosts with a zero-data-retention policy.
  • Nothing you upload is used to train any model — not by Nia, and not by the services named below, under their published policies.
  • Usage is capped. When your plan's monthly allowance is reached, AI work stops rather than continuing at cost.
  • No decision about you is made solely by a machine.

1. What happens to a document you upload

Excel and CSV files are read by the platform's own, rule-based extractor. The file is not sent to any AI service to be read. If a sheet yields nothing, the platform may send a bounded portion of that sheet's text (at most 12,000 characters per sheet, a limited number of sheets, within a time budget) to an AI model to find the indicators it asks for.

Word files are read on the platform first. A Word file with no readable text (for example, a document made entirely of images) is treated like a scan, and a Word or PDF data submission in which the platform's own reader finds no reportable figures is also sent to the parsing service so the figures can be found in the narrative.

PDFs, images and scans cannot be read reliably by rules alone. The whole file is sent to an external document-parsing service, LlamaParse by LlamaIndex, which returns text and tables. The platform sends every job with the "do not cache" instruction, and LlamaIndex's published policy is that files are used only to return your result, never for model training. LlamaIndex offers processing in Europe as well as the United States; the platform uses the European region, so files are processed in the European Union.

For investee and fund submissions the order is strict: the platform's rule-based reader goes first, and a file it can read is never sent anywhere. Only a scan, a photo, a PDF whose columns cannot be told apart, or a prose-only Word document goes to the parsing service.

After parsing, the platform maps the returned tables to indicators using rules, not AI. If that produces nothing, the document's text may be sent to an AI model to extract indicators from prose.

The platform keeps a copy of the parsed text and tables, keyed to your organisation, so that re-uploading an identical file does not need a second parse.

2. Figures reconstructed by AI are flagged, and must be acknowledged

When a submission's figures were reconstructed from a scan or prose rather than read from cells, the platform marks the whole submission as AI-reconstructed. The mark is written into the data-quality record, shown on the review screen, and the submission cannot be approved until a named person records a written reason for accepting it. The same rule applies to fund-level data. The flag is recorded per file, not per individual figure.

3. The Advisor and Policy Review

The Advisor answers only from material it has retrieved: Nia's knowledge base and, if you have uploaded them, your organisation's own documents. It has no web access and is instructed not to use outside knowledge. What is sent to the model for a question is: your message, the passages retrieved for it, earlier turns of the conversation, any document you attached to that conversation (in full), and — when you ask about your portfolio — a read-only summary of your own submission status, indicator trends and data-quality counts, computed by the platform. Retrieval is confined to your organisation's data.

Documents you add to the Policy Hub or attach to a conversation are read on the platform and split into passages; each passage is sent to an embedding model so it can be found again. Embedding hosts are limited to two named providers that do not store or train on the input. Policy Review sends the policy's text to an AI model to compare it against Nia's best-practice guidance; the result is a report for you to read, not a change to your documents.

Conversation history is kept so you can return to it. The platform's retention setting for conversations is 12 months. The automated purge that enforces it is a planned control that is not yet in operation; until it runs, a conversation is kept until you delete it.

4. Where AI text goes — the model hosts

All AI model calls (other than document parsing) are routed through OpenRouter. Every request carries three restrictions written into the platform:

  1. A fixed list of allowed hosts — Anthropic, Google Vertex AI, Amazon Bedrock, Microsoft Azure and OpenAI. A request that cannot be served by one of them fails; it is never sent elsewhere.
  2. "Deny data collection" — no host that stores prompts or trains on them may be used.
  3. Zero data retention for conversational and extraction models — only endpoints with a published zero-retention policy. (Embedding models have no zero-retention endpoint, so for them restrictions 1 and 2 apply.)

Today the Claude models are served on Amazon Bedrock and Google Vertex AI zero-retention endpoints, and embeddings on Microsoft Azure or OpenAI; Anthropic's own service is a permitted host but receives no requests, because no zero-retention endpoint is listed there for the models the platform uses.

OpenRouter's own policy is that it does not retain prompts unless the account opts in to logging. [to be confirmed: prompt logging is off on the OpenRouter account the platform uses.]

Models in use today are Anthropic's Claude models (Claude Haiku 4.5 for extraction, quality checks, the Advisor and guidance drafting; Claude Sonnet 4 for report narratives) and an OpenAI embedding model (text-embedding-3-small). Nia can switch a task to a Google Gemini or OpenAI GPT model only from a short vetted list; this page is updated if that happens.

5. What never happens

  • Your files, figures and conversations are not used to train models — the platform sends nothing for training, and the named services state that they do not train on customer data.
  • No file that the platform can read itself is sent to an external service.
  • No AI output is written into your submissions, reports, indicator dictionary or action tracker without a person approving it. (Advisor answers are stored in your conversation history so you can return to them; they do not change any of that data.)
  • No decision with legal or similarly significant effect on anyone is made solely by automated processing. AI on this platform assists with reading documents, drafting and answering questions; approvals, submissions and reports remain human decisions.

6. Usage limits

Each metered AI operation reserves one "run" from your plan's monthly allowance before it starts. At the cap, the operation is refused and nothing is sent to a model. See Usage Caps for what counts as a run.

7. Your responsibility

AI output on this platform is assistance. The fund remains responsible for the data it submits, approves and reports. Check reconstructed figures against the source document, read Advisor answers with their citations, and record your reason when you accept an AI-reconstructed submission — that reason is part of the audit record.

8. Who processes your data

The services named on this page — LlamaIndex (document parsing), OpenRouter and the five model hosts above (AI models and embeddings) — act as processors for your organisation's content. The full list of sub-processors, with regions and safeguards, is published on our sub-processor list and summarised in the Privacy Policy. [to be confirmed: processing agreements accepted or signed with each provider] [under legal review]

On this page