AppFlowy-IO / AppFlowy-IO/AppFlowy

[FR] Add granular control to exclude pages/blocks from embeddings (like .gitignore for AI indexing)

Abierto
#8,701 2 comentarios 0 reacciones 0 asignados Ver en GitHub
AI new feature self-hosted
Lenguaje dominante
Dart
Estrellas
76.6k
Forks
6k
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

### Description

AppFlowy currently automatically **sends all workspace content to the configured embeddings provider** to enable semantic search and AI-powered features.

However, there is no mechanism to exclude specific pages, documents, or blocks from being included in the embedding/indexing process.

This limits control over what data is processed by AI systems, especially for users integrating AppFlowy with external or self-hosted AI pipelines (e.g. local LLM servers, vector databases, or custom memory systems).

### 🛠️ Suggested implementation approach

This feature could be implemented in a lightweight and extensible way without changing the core embedding pipeline:

### 1. Metadata-based exclusion (recommended baseline)
Allow pages or blocks to include a simple flag:

`ai_index: false`

or

`embedding: disabled`

This would allow the embedding pipeline to skip content during chunking.

### 2. Tag-based exclusion rules
Support tags such as:
- no-ai
- private
- no-index

which would automatically exclude tagged content from embedding generation.

### 3. Optional .aiignore file (workspace-level rules
Similar to .gitignore, allowing patterns such as:
`/private/*
/journal/drafts/*
*.secret`

This would enable flexible workspace-wide exclusion rules.

4. Hierarchical policy resolution
In case of conflicts, rules could follow a priority system:

`Block-level rule > Page-level rule > Workspace-level rule`

### Impact

The absence of granular embedding control introduces architectural and usability constraints for advanced and privacy-sensitive use cases:

- Global indexing without exclusion control: **all workspace content is embedded by default,** with no ability to opt out specific pages or blocks
- Lack of data boundary separation: **no distinction** between AI-searchable content and private or non-indexed information
- Privacy and governance limitations: users cannot enforce selective embedding policies for sensitive or internal data
- Reduced compatibility with external AI architectures: limits integration with custom memory systems, vector databases, and agent-based workflows
- Scalability constraints for advanced setups: makes it difficult to adopt AppFlowy in production-grade or hybrid AI environments where data control is required

### 💡 Why this matters

For modern AI-driven workflows, embedding pipelines are often treated as a controlled data layer rather than a global automatic process.

🚨 Without granular control, AppFlowy behaves as a fully inclusive indexing system, which reduces flexibility for advanced users and self-hosted deployments.

### Additional Context

_No response_

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.