Skip to content

Extraction Pipeline

FluidRAG can extract Freshdesk support tickets into your knowledge graph on a schedule, making customer issue history available for RAG queries alongside your other ingested documents.

The pipeline runs in three stages:

  1. Ticket Extraction — FluidRAG fetches tickets from Freshdesk (with descriptions and conversations included), paginating 100 per page. Internal notes are split from public conversations, and agent/requester IDs are resolved to display names. Each ticket is saved as a standalone Markdown file at /app/extraction_data/freshdesk/raw/ticket_{id}.md.

  2. LLM Preprocessing — Each raw ticket is analyzed by an LLM to determine whether it contains viable knowledge for a company-wide RAG system. Viable tickets are rewritten into a structured document with # Topic, # Summary, and # Detailed Information sections, validated for completeness, and written to /app/extraction_data/freshdesk/processed/[{date}]ticket_{id}_{title}.md. Non-viable tickets (e.g., no technical content) are skipped.

  3. Knowledge Graph Ingestion — Processed Markdown documents are ingested into FalkorDB with entity extraction, embedding generation, and relationship creation, making them searchable via graph and vector search.

Each raw Markdown file contains frontmatter metadata followed by the ticket body:

Frontmatter FieldDescription
sourcefreshdesk
source_urlLink to the ticket in Freshdesk
source_id / ticket_idFreshdesk ticket ID
titleTicket subject
created_at / updated_atISO 8601 timestamps (UTC)
requester / responderResolved display names
tenant_namecf_portal_name_ak custom field
priority / service_impact / resolution_detailRelevant custom field values
statusTicket status code

The body includes the ticket Description, a Custom Fields section (with readable labels), Conversations (public), and Internal Notes (private). HTML is cleaned to plain text.

  1. Open the Admin UI at http://localhost/config-ui
  2. Navigate to the Connectors page
  3. Find Freshdesk in the connector list
  4. Toggle Enabled to on
SettingDefaultDescription
freshdesk_enabledtrueEnable the Freshdesk extraction source (hot-reload)
freshdesk_domain(empty)Freshdesk subdomain (e.g., company.freshdesk.com)
ActionHow
Trigger extraction nowConnectors → Freshdesk → Trigger Extraction (always runs a forced sync)
Reset sync stateConnectors → Freshdesk → Reset Sync State — clears the last sync timestamp so the next run extracts everything
Configure scheduleSet a cron expression (e.g., 0 */8 * * *) in the Freshdesk connector’s extraction settings

All endpoints require admin authentication.

EndpointMethodDescription
/api/connectors/freshdeskGETGet Freshdesk connector status and configuration
/api/connectors/freshdeskPATCHEnable/disable the Freshdesk connector (hot-reload)
/api/connectors/freshdesk/extraction/triggerPOSTTrigger extraction immediately (forced sync; optional stages body parameter)
/api/connectors/freshdesk/extraction/statusGETGet Freshdesk extraction status
/api/connectors/freshdesk/extraction/historyGETGet Freshdesk extraction history
/api/connectors/freshdesk/extraction/resetPOSTReset sync state for full resync
/api/connectors/freshdesk/source-configGETGet Freshdesk source configuration
/api/connectors/freshdesk/source-configPATCHUpdate Freshdesk source settings
/api/connectors/freshdesk/schedulePATCHSet the extraction schedule (cron expression)
/api/connectors/freshdesk/scheduleDELETERemove the extraction schedule