---
title: "创建知识库"
method: POST
path: "/knowledge-bases"
tags: ["知识库"]
---

# 创建知识库

`POST /knowledge-bases`

创建新的知识库

## Request body

- GithubComTencentWeKnoraInternalTypesKnowledgeBase
  - `asr_config` GithubComTencentWeKnoraInternalTypesASRConfig
    - `enabled` boolean
    - `language` string — optional: language hint for transcription
    - `model_id` string
  - `auto_tag_config` GithubComTencentWeKnoraInternalTypesAutoTagConfig
    - `enabled` boolean
    - `max_tags` integer
    - `model_id` string
    - `skip_if_tagged` boolean — SkipIfTagged leaves documents that already carry tags untouched, so a deliberate manual classification is not diluted by model guesses. It is a pointer because the default is true: rows written before this field existed decode to nil and must not silently flip to "always append".
  - `chunk_count` integer — Chunk count (not stored in database, calculated on query)
  - `chunking_config` GithubComTencentWeKnoraInternalTypesChunkingConfig
    - `child_chunk_size` integer — ChildChunkSize is the size of child chunks used for embedding (default: 384). Only used when EnableParentChild is true.
    - `chunk_overlap` integer — Chunk overlap
    - `chunk_size` integer — Chunk size
    - `enable_parent_child` boolean — EnableParentChild enables two-level parent-child chunking strategy. When enabled, large parent chunks provide context while small child chunks are used for vector matching. Retrieval matches on child but returns parent content.
    - `languages` string[] — Languages hints the heuristic patterns. Empty = auto-detect from content. Examples: ["de"], ["en", "zh"].
    - `parent_chunk_size` integer — ParentChunkSize is the size of parent chunks (default: 4096). Only used when EnableParentChild is true.
    - `parser_engine_rules` GithubComTencentWeKnoraInternalTypesParserEngineRule[] — ParserEngineRules configures which parser engine to use for each file type. When empty, the builtin engine is used for all types.
      - `engine` string
      - `file_types` string[]
      - `xlsx_first_row_as_header` boolean — XLSXFirstRowAsHeader restores row-1 column context for flat XLSX tables. nil preserves the parser default; an explicit false disables the mode.
    - `separators` string[] — Separators
    - `strategy` string — Strategy selects the adaptive chunking tier. Empty / "legacy" preserves the historical recursive splitter; "auto" lets a profiler pick between heading-aware, heuristic and recursive tiers; "heading" / "heuristic" / "recursive" pin the tier explicitly.
    - `table_metadata_instructions` string — TableMetadataInstructions contains optional business guidance used when generating searchable summaries for CSV/Excel tables. The system-owned output contract remains fixed; these instructions only add domain context.
    - `token_limit` integer — TokenLimit caps chunk size in approximate tokens. 0 = use ChunkSize as a character count.
  - `created_at` string — Creation time of the knowledge base
  - `creator_id` string — CreatorID records the user ID of whoever originally created the KB. Used by the workspace-level RBAC middleware to let Contributors edit their own KBs without granting them access to everyone else's. Nullable for backward compatibility with rows created before the RBAC migration backfilled the column to the workspace Owner.
  - `creator_name` string — CreatorName 是 CreatorID 对应用户的展示名（username / email 等）， 仅在列表场景由 handler 批量回填，不落库；为空表示创建者无法解析（用户已删除、 CreatorID 为空的老数据等）。前端用它在卡片来源徽章上做 mine vs workspace 的二分。
  - `deleted_at` GormDeletedAt
    - `time` string
    - `valid` boolean — Valid is true if Time is not NULL
  - `description` string — Description of the knowledge base
  - `embedding_model_id` string — ID of the embedding model
  - `extract_config` GithubComTencentWeKnoraInternalTypesExtractConfig
    - `custom_instructions` string — CustomInstructions adds domain-specific extraction guidance while the system keeps ownership of the structured graph output protocol.
    - `enabled` boolean
    - `nodes` GithubComTencentWeKnoraInternalTypesGraphNode[]
      - `attributes` string[]
      - `chunks` string[]
      - `name` string
    - `relations` GithubComTencentWeKnoraInternalTypesGraphRelation[]
      - `node1` string
      - `node2` string
      - `type` string
    - `tags` string[]
    - `text` string
  - `faq_config` GithubComTencentWeKnoraInternalTypesFAQConfig
    - `index_mode` 'question_only' | 'question_answer'
    - `question_index_mode` 'combined' | 'separate'
  - `id` string — Unique identifier of the knowledge base
  - `image_processing_config` GithubComTencentWeKnoraInternalTypesImageProcessingConfig
    - `model_id` string — Model ID
  - `indexing_strategy` GithubComTencentWeKnoraInternalTypesIndexingStrategy
    - `graph_enabled` boolean — GraphEnabled enables knowledge graph entity/relation extraction
    - `keyword_enabled` boolean — KeywordEnabled enables keyword-based (BM25) search
    - `vector_enabled` boolean — VectorEnabled enables semantic vector embedding and search
    - `wiki_enabled` boolean — WikiEnabled enables automatic wiki page generation from documents
  - `is_pinned` boolean — IsPinned and PinnedAt are computed per-caller from user_kb_pins (see migration 000050). They used to be stored on the row itself, which made pinning a workspace-wide ordering decision gated behind the kb-edit RBAC guard. The columns are still present in legacy schemas for rollback safety but are no longer read or written by the application — both fields are tagged `gorm:"-"` so GORM ignores them on every CRUD call and the list handler stamps them after enriching with the caller's pin set.
  - `is_processing` boolean — IsProcessing indicates if there is a processing import task (for FAQ type knowledge bases)
  - `is_temporary` boolean — Whether this knowledge base is temporary (ephemeral) and should be hidden from UI
  - `knowledge_count` integer — Knowledge count (not stored in database, calculated on query)
  - `name` string — Name of the knowledge base
  - `pinned_at` string — PinnedAt records when the current caller pinned this knowledge base; nil when they have not.
  - `processing_count` integer — ProcessingCount indicates the number of knowledge items being processed (for document type knowledge bases)
  - `question_generation_config` GithubComTencentWeKnoraInternalTypesQuestionGenerationConfig
    - `custom_instructions` string — CustomInstructions describes the intended audience or question style. It is appended to the stable system question-generation template.
    - `enabled` boolean
    - `question_count` integer — Number of questions to generate per chunk (default: 3, max: 10)
  - `share_count` integer — ShareCount indicates the number of organizations this knowledge base is shared with (not stored in database)
  - `storage_backend_id` string — StorageBackendID binds this KB to one concrete storage instance. The legacy provider field remains readable during migration only.
  - `storage_config` GithubComTencentWeKnoraInternalTypesStorageConfig
    - `app_id` string — App ID (COS specific)
    - `bucket_name` string — Bucket Name
    - `endpoint` string — Endpoint (S3 specific) - e.g., s3.amazonaws.com, oss-cn-hangzhou.aliyuncs.com
    - `force_path_style` boolean — ForcePathStyle (S3 specific) - whether to use path-style URLs
    - `path_prefix` string — Path Prefix
    - `provider` string — Provider: "cos", "minio", "s3"
    - `region` string — Region
    - `secret_id` string — Secret ID (COS) / Access Key ID (S3, MinIO)
    - `secret_key` string — Secret Key (COS) / Secret Access Key (S3, MinIO)
    - `use_ssl` boolean — UseSSL (S3 specific) - whether to use HTTPS
  - `storage_provider_config` GithubComTencentWeKnoraInternalTypesStorageProviderConfig
    - `provider` string — "local", "minio", "cos", "tos", "s3", "oss", "ks3", "obs"
  - `summary_model_id` string — Summary model ID
  - `tenant_id` integer — Workspace ID
  - `type` string — Type of the knowledge base (document, faq, etc.)
  - `updated_at` string — Last updated time of the knowledge base
  - `vector_store_id` string — VectorStoreID references the VectorStore this knowledge base is bound to. When nil, the KB falls back to the workspace's effective engines derived from the RETRIEVE_DRIVER environment variable (env store flow). This field is set once at creation time and must not be modified afterwards; enforcement lives at the GORM layer (`<-:create`) plus the service-layer KB update path, which omits this field from its update DTO.
  - `vlm_config` GithubComTencentWeKnoraInternalTypesVLMConfig
    - `api_key` string — API Key
    - `base_url` string — Base URL
    - `custom_instructions` string — CustomInstructions adds KB-specific image interpretation guidance without replacing the system-owned OCR and Markdown output contract.
    - `description_language` string — DescriptionLanguage controls the language used for generated image captions. Empty means follow the document/request language.
    - `enabled` boolean
    - `interface_type` string — Interface Type: "ollama" or "openai"
    - `model_id` string
    - `model_name` string — 兼容老版本 Model Name
  - `wiki_config` GithubComTencentWeKnoraInternalTypesWikiConfig
    - `content_instructions` string — ContentInstructions controls tone, structure and emphasis for generated summary/entity/index prose. Citation and merge rules remain system-owned.
    - `extraction_granularity` 'focused' | 'standard' | 'exhaustive'
    - `extraction_instructions` string — ExtractionInstructions tells candidate extraction which domain concepts to emphasize without replacing the stable JSON/citation protocol.
    - `ingest_batch_size` integer — IngestBatchSize controls how many pending ops a single batch claims and processes before scheduling a follow-up. 0 falls back to the hard-coded default (5). Larger batches amortize per-batch setup and let more docs share the batch-internal Map/Reduce fan-out; smaller batches spread a KB's backlog across more concurrent batches (finer scheduling grain).
    - `ingest_map_parallel` integer — IngestMapParallel sets the errgroup limit for the Map phase (per-document extraction + summary + chunk citation) WITHIN one batch. 0 falls back to 10. Bound by the LLM provider's concurrency limit and the worker's outbound HTTP pool. Remember it multiplies with the number of concurrent batches (IngestMaxInflight).
    - `ingest_max_inflight` integer — IngestMaxInflight caps how many ingest batches for THIS KB may run concurrently in the shared wiki worker pool (standard/Redis mode only). 0 falls back to the hard-coded default (4). Since Phase 3 removed the exclusive per-KB lock, one KB's backlog could otherwise monopolize the whole pool during a bulk import and starve other KBs; this knob trades a single KB's peak throughput for cross-KB fairness. Set it >= the wiki pool size to effectively disable the cap.
    - `ingest_reduce_parallel` integer — IngestReduceParallel sets the errgroup limit for the Reduce phase (per-slug page write) WITHIN one batch. 0 falls back to 10. Bound by the same LLM concurrency / HTTP pool considerations as the Map phase, plus DB connection pool size. Same multiplier caveat as IngestMapParallel.
    - `max_pages_per_ingest` integer — MaxPagesPerIngest limits pages created/updated per ingest operation (0 = no limit)
    - `synthesis_model_id` string — SynthesisModelID is the LLM model ID used for wiki page generation and updates

## Response `201`

创建的知识库

- object

## Other responses

- `400` — 请求参数错误

---

[API](https://skmtc.net/tencentblueking/apis/weknora-api.md) · [All operations](https://skmtc.net/tencentblueking/apis/weknora-api/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/tencentblueking/weknora-api/revisions/abce9036def5/schema)
