Most enterprise generative-AI programmes started with a demo that impressed the executive committee and then quietly died. The exceptions share a pattern: they stopped trying to build "an AI assistant" and built something much narrower — a document-grounded chatbot that answers questions about internal knowledge, inside the tool people already have open all day. That is Slack or Microsoft Teams, and the architecture behind it is retrieval-augmented generation (RAG).
This article covers the use cases that actually reach production, how to wire a RAG assistant into Slack and Teams properly, how to handle the access-control problem that kills most pilots, and how to compute a ROI you can defend in front of a CFO.
Which use cases survive contact with production
A knowledge-base chatbot works when three conditions are met: the questions are repetitive, the answers exist somewhere in writing, and being wrong is annoying rather than catastrophic. Filter your backlog through that lens and a short list emerges.
- Internal IT / tier-1 support. "How do I request a VPN token?", "What's the procedure for a lost laptop?". High volume, documented answers, measurable ticket deflection.
- HR and people operations. Leave policies, expense rules, mobility, benefits. Same profile, and it relieves a team that is chronically interrupted.
- Pre-sales and bid support. Answering RFP questions from past proposals, security questionnaires, reference architectures. Here the value is cycle time, not headcount.
- Engineering onboarding and runbooks. Grounded on the internal wiki, ADRs, Terraform modules and incident post-mortems. This one gets the highest organic adoption because engineers live in Slack.
- Quality, compliance and regulated procedures. ISO, GxP, internal control. Answers must carry citations — which is precisely what RAG gives you and a fine-tuned model does not.
Conversely, be sceptical of "ask anything about the company" scopes, of chatbots expected to compute figures from spreadsheets, and of external customer-facing bots as a first project. The failure cost is far higher outside the firewall.
RAG, fine-tuning, or just a huge context window?
The question comes up in every steering committee. The honest answer is that they solve different problems.
| Approach | Good at | Weak at | Ops cost |
|---|---|---|---|
| RAG | Fresh, citable, permission-filtered answers over a corpus that changes weekly | Complex multi-hop reasoning across many documents | Ongoing: ingestion pipeline, index, evaluation |
| Fine-tuning | Tone, output format, domain vocabulary, structured extraction | Factual freshness, citations, per-user access control | Re-training on every knowledge change — untenable for a wiki |
| Long context (full docs in the prompt) | Small, stable corpora; deep analysis of one document | Cost and latency scale with corpus size; retrieval quality degrades in the middle of very long contexts | Low to build, high per query |
For a knowledge-base assistant, RAG is the default. Long context is a useful complement once the retriever has narrowed things down: retrieve broadly, then feed the model whole sections rather than 300-token fragments.
Anatomy of a RAG pipeline that holds up
The demo version — embed everything, cosine similarity, top-5, prompt — reaches perhaps 60% useful answers, which is below the trust threshold at which users stop coming back. The production version adds four things.
Structure-aware chunking. Split on headings, not on a fixed character count. Carry the document title and heading path into each chunk's text so that an isolated paragraph still says what it is about. Keep tables intact; a chunk that cuts a table in half is worse than no chunk.
Hybrid retrieval. Dense vectors miss exact identifiers — error codes, product SKUs, internal acronyms. Combine BM25 and vector search and fuse the two ranked lists (reciprocal rank fusion works well and needs no tuning). This is usually the single biggest quality jump.
Re-ranking. Retrieve 30–50 candidates, then score them with a cross-encoder reranker and keep the best 5–8. It costs a few tens of milliseconds and dramatically reduces the amount of irrelevant context reaching the model.
Mandatory citations and abstention. The system prompt must instruct the model to answer only from the provided passages, cite the source of each claim, and explicitly say it does not know otherwise. A chatbot that says "I couldn't find this in the documentation" earns more trust than one that improvises.
The problem nobody plans for: access control
Your wiki, your drive and your ticketing system all have permissions. Your vector index, by default, does not. Ship without solving this and the first time the bot quotes a compensation review to an intern, the project is over.
The workable pattern is to propagate ACLs at ingestion time into the vector payload, then filter at query time with the requesting user's identity — resolved from Slack or Entra ID, not from anything the user typed.
from qdrant_client.models import Filter, FieldCondition, MatchAny
# groups resolved server-side from the IdP, never from the chat message
user_groups = directory.groups_for(slack_user_id)
acl = Filter(must=[FieldCondition(
key="acl_groups",
match=MatchAny(any=user_groups + ["all-employees"]),
)])
hits = qdrant.query_points(
collection_name="kb",
query=embedding,
query_filter=acl,
limit=40,
)Two operational consequences. First, re-index when permissions change, not only when content changes — a document moved to a restricted space must disappear from the index within minutes. Second, log every answer with the user, the retrieved chunk IDs and the citations. You will need that trail for the security review, and for debugging.
Slack vs Teams: not the same integration project
The retrieval backend is identical; the surface layer is not.
| Slack | Microsoft Teams | |
|---|---|---|
| SDK | Bolt (Python/JS), Events API, Socket Mode for dev | Bot Framework SDK + Azure Bot Service, or Teams AI Library |
| Identity | Slack user ID → mapped to your IdP | Entra ID natively, SSO on-behalf-of flows available |
| UX primitives | Threads, Block Kit, slash commands, Assistant panel | Adaptive Cards, message extensions, tabs, meeting context |
| Time to first version | Fast — a day for a working prototype | Slower — app manifest, Azure resources, tenant admin approval |
| Main friction | 3-second ack window, needs async responses | Governance and publishing workflow through IT |
from slack_bolt import App
app = App(token=os.environ["SLACK_BOT_TOKEN"])
@app.event("app_mention")
def on_mention(event, client, ack):
ack() # respond within 3s, then work asynchronously
thread = event.get("thread_ts", event["ts"])
placeholder = client.chat_postMessage(
channel=event["channel"], thread_ts=thread, text="Searching the knowledge base…"
)
result = rag.answer(
question=strip_mention(event["text"]),
user_id=event["user"],
history=load_thread(client, event["channel"], thread),
)
client.chat_update(
channel=event["channel"], ts=placeholder["ts"],
blocks=render_answer_with_citations(result),
)Whichever platform you pick: always reply in a thread, always render citations as clickable links to the source document, and always attach 👍/👎 feedback actions. That feedback is your only cheap source of evaluation data.
Computing a ROI you can defend
Avoid the vendor slide that multiplies headcount by a productivity percentage. Build the number bottom-up, and measure a baseline before you launch.
Monthly gross gain =
Q (questions handled by the bot per month)
× d (deflection rate: share genuinely resolved without a human)
× t (average human time saved per question, in hours)
× c (fully loaded hourly cost of the person who would have answered)
Monthly net gain = gross gain − (inference + vector DB + hosting
+ ingestion pipeline run
+ amortised build + product ownership)Three disciplines make this credible. Measure Q and t from the existing system: ticket counts by category before go-live, and the median resolution time of the tickets you intend to deflect. Do not assume d: derive it from thumbs-up rates and from the drop in tickets in the targeted categories, and expect it to be well below your optimistic estimate at first. Count the run cost honestly: inference tokens are usually not the dominant line. A part-time product owner curating content and reviewing failed answers is, and it is the line that actually determines success.
The secondary benefits are real but harder to book: faster onboarding, fewer interruptions for senior experts, and — underrated — the fact that a chatbot surfaces exactly which documentation is missing or contradictory. Several teams get more value from that diagnosis than from the answers themselves.
Evaluation and observability
Treat the assistant like any other production service, with two extra layers. On the retrieval side, keep a golden set of 100–200 real questions with known correct source documents, and track recall@k in CI whenever chunking, embeddings or the reranker change. On the generation side, check groundedness — is every claim supported by a retrieved passage — using an LLM-as-judge on a sample, plus human review of all thumbs-down.
Instrument the usual suspects too: end-to-end latency percentiles, token cost per conversation, abstention rate, share of questions with zero retrieved documents above threshold (your content gap indicator), and weekly active users per team. Adoption curves per team tell you more about value than any aggregate satisfaction score.
A realistic 90-day plan
Weeks 1–3: pick one domain and one owner, inventory the sources, and collect 150 real questions with their expected answers. Build the ingestion pipeline and the hybrid index. Weeks 4–7: RAG API with ACL filtering, reranking, citations and abstention; iterate against the golden set until retrieval recall is clearly above what a keyword search returns. Weeks 8–10: Slack or Teams integration, feedback buttons, full logging, security and DPO review. Weeks 11–13: pilot with 30–80 users, weekly triage of failures, fix the documentation the bot exposes as missing, then decide on extension with real numbers in hand.
One last point that matters for European organisations: nothing in this architecture forces you to send your internal knowledge outside the EU. Open-weight or European models served on EU-hosted GPU infrastructure, a self-hosted vector database, and an ingestion pipeline inside your own VPC produce a system whose quality is driven mainly by retrieval — not by having the largest possible model. Start with the retrieval, and the model choice becomes a swappable detail.
