Taqania

Public Sector · Applied AI

Building an intelligent retrieval engine for a two-decade document archive

Intelligent retrieval with mandatory citation restricted by user permissions, ensuring search engines do not become an information leak vulnerability.

Illustrative examples simulating real-world scenarios; designed to showcase engineering methodology, analysis style, and target outcomes with precision and transparency.

What we would target

11 weeksTarget timeline from kickoff to production
20 yearsScope of archived documents and correspondence
480Approved, human-labeled evaluation questions

The problem

Two decades of scanned documents and policies accumulated in a government entity; making information access impossible except for those who know exactly where it is, crippling institutional memory.

What we would do

  • Processing Arabic and scanned documents as a core requirement, validating quality before committing to any AI model.
  • Inheriting access permissions from existing accounts to prevent any document from surfacing to an unauthorized user.
  • Forcing the model to cite the original document and paragraph, and applying an accuracy gate measuring performance against 480 real, human-labeled queries.

The point of it

The real challenge in retrieval systems lies in unifying the source of truth and enforcing permissions, not simply in invoking AI models.

Tell us what you are trying to solve.

Share the operational challenge and the regulatory or technical constraints you are working within. A consultant replies within two working days, in Arabic or English.