Search across all content
Adds semantic search, trend visualisation, and automated summaries across millions of records.
Public sector organisations generate an immense, continuously expanding volume of text-based information across multiple layers of operation. For years, the City of Edmonton faced severe operational friction due to high-value data being trapped inside disparate corporate silos. Critical organisational resources—ranging from Council report files and media releases to historic survey responses, bylaws, and memos—were stored independently across different software tools and web pages.
The baseline status quo forced policy analysts, city planners, and administrative leaders to spend an excessive amount of time manually sorting through separate systems like Google or the internal corporate intranet to track down historical precedents or public feedback. Under financial constraints that barred the procurement of expensive proprietary enterprise search licenses, the city needed a method to consolidate millions of records. The core structural challenge was to design a platform capable of handling unstructured text documents while maintaining strict data governance, local context relevance, and zero licensing fees.
The Data Science and Research team engineered and scaled Text Depot—a centralised, context-aware semantic search platform tailored specifically for internal municipal operations.
Part 1: Technology Stack and Open Source Infrastructure
To ensure cost efficiency, the platform was constructed entirely from open-source technologies:
Part 2: Implementation Mechanics, AI Annotations, and Security
To transform raw documents into searchable intelligence, the team built a uniform ingestion pipeline where every record is stored in a standardised format, including the full document text, source titles, dates, parent document items (where available), and URLs. As data flows into the Elasticsearch index, the system attaches automated metadata layer annotations using advanced NLP models:

Text Depot successfully converted isolated data sources into an accessible centralised repository for city staff and senior administration.
1. Massive Scale Integration
Centralised over 1.4 million corporate data records, including approximately 40,000 Council reports, 30,000 daily media monitoring articles, thousands of media releases, active bylaws, carried Council motions, and historic engagement surveys.
2. Widespread User Adoption
Scaled to whole-organisation utility, reaching an active milestone of 1,373 unique internal users and logging over 9,500 total user sessions to date—with more than 5,500 sessions occurring since 2024 alone.
3. Informed Governance & Safeguards
Adopted directly by executive leadership and City Councillors to navigate contextual histories instantly. The internal tool architecture also successfully caught public-facing data leaks (re-securing legacy public tools like eScribe when sensitive fields inadvertently surfaced).
Prioritise Internal Product Stewardship and Strict Governance. When handling data containing potentially identifiable public notes or historic surveys, organisations must institute granular, role-based access restrictions that undergo formal evaluation on an annual basis to maintain strict privacy compliance.
Maximise Capital Efficiency with Built-in Tool Ecosystems. Edmonton demonstrated that an elite digital environment can be established using a completely open-source, non-proprietary tech stack (Elasticsearch, R/Shiny, Docker Swarm) without buying costly commercial end-to-end software suite licenses. Leveraging open-source code allows internal development teams to stay agile and continuously update analytical functions dynamically.
Localise NLP Models to Understand Community Semantics. Standard, out-of-the-box AI tools fail to capture local institutional memory and context. Supplementing generic vector databases with custom models trained directly on specialised local vocabularies guarantees that regional concepts, neighbourhood groupings, and idiosyncratic civic terminology are successfully parsed to return hyper-accurate results.
Launch year: 2021





Connect with 500,000+ public servants solving your hardest challenges.





Connect with 500,000+ public servants solving your hardest challenges.
Help public servants worldwide learn from your work, what worked, what flopped and what you'd do differently
Share your project
Log in or sign up to continue the conversation