# INITKOA CONTEXT PACK repository: Rejean-McCormick/SemantiK_Architect source_commit: cf1cf76cc2a95c50b933d49e1afaff7c28117ffd source_mode: git working_tree_markdown: clean working_tree_selected: clean selection_mode: markdown wiki_source_commit: cbc326182d54c35df74b244751714d19aa2e7869 wiki_working_tree_markdown: clean policy_version: 2026-09-10.13 repo_files: 42 wiki_files: 25 source_files: 67 included_files: 67 excluded_files: 0 duplicate_files: 0 content_bytes: 403932 authority_counts: {"reference":67} content_role_counts: {"knowledge":42,"navigation":25} generated_at: 2026-09-10T13:04:11-04:00 files: 67 content_sha256: bf84f1ad2bd52825309b6dff457b75401dbd5208cd7ec5fd88f22d3442363c8e ================================================================================================ FILE INDEX ================================================================================================ 01. [reference] [navigation] wiki/_Footer.md | bytes=267 | sha256=6393d670aaed7c8fe1b6a7caf54ff6b8d0e2fdb4e5a20d288b02684da33109cf 02. [reference] [navigation] wiki/_Sidebar.md | bytes=1035 | sha256=76b6cb73b2a55f1d0747a204525d62faab72babeb56684bce8da8d442baaef85 03. [reference] [navigation] wiki/Add-a-Language-High-Level.md | bytes=3826 | sha256=ac05585c7caf28b5f2e978379caa39cf30f7aa00256ec2a8c7955b67ed174064 04. [reference] [navigation] wiki/API-Overview.md | bytes=2275 | sha256=e161a286f53d2a1a8a2aea8c462d3bb4554a3ee853b61c375a39a08a4e375286 05. [reference] [navigation] wiki/Changelog.md | bytes=1973 | sha256=faf9abec1521092ff6cfdb5cec7bc250142067733c1c285e63fa84bd3eb38245 06. [reference] [navigation] wiki/Conceptual-Flow-Meaning-to-Text.md | bytes=2456 | sha256=b584eb543da5bcb0216ea520f937fd0cba9410d384abaa78d42ec6d7b31310a9 07. [reference] [navigation] wiki/Context.md | bytes=2462 | sha256=b43ca7a23af89688ad77573eac023612be4546a836dc1b92862d1f627fd4b193 08. [reference] [navigation] wiki/Correctness-and-Verifiability-Gold-Standards-UD-Export-Judge.md | bytes=2639 | sha256=17c7a2601aaed1482f76128bd9797c84655dfd2d96633778fb6b3235a38c07bd 09. [reference] [navigation] wiki/Decisions.md | bytes=2718 | sha256=9cc35f8c8294ca251148436b3f56a4438f91580c047a4d0594c0acdd82f38f12 10. [reference] [navigation] wiki/Glossary.md | bytes=5599 | sha256=484745566f66e08a76ae0d4060f0e02a367b10996863b294aa7927fe88307403 11. [reference] [navigation] wiki/Grammar.md | bytes=3167 | sha256=d06b24c10f0943488dee118bc721e6f7e05112ca03602208b78a205be595aaa4 12. [reference] [navigation] wiki/Home.md | bytes=3567 | sha256=2d8f622c87e1d03b67492b5505007dde746ba652ed18610ebee763157403e4fe 13. [reference] [navigation] wiki/Improve-a-Language-High-Level.md | bytes=5474 | sha256=0647d1388d6e6c3aeac503a21c7fb2b3258abc622ffe2f26dbe0dd6ce8f3bdc7 14. [reference] [navigation] wiki/Inputs-Frames.md | bytes=1783 | sha256=776f70ee6a76ab8edf7bd2ba26631addf2c7974f761e989d08da1b41eac6b9b4 15. [reference] [navigation] wiki/Inputs-Ninai.md | bytes=1754 | sha256=bfba15dabfe14fc6aad44a8c20af8842808b4067249e73ea8672588b4c146b29 16. [reference] [navigation] wiki/Language-Coverage-Strategy-Tiers.md | bytes=1723 | sha256=03ef5d2f259614646500810b5b347d496391d5f7e5dae8d64b55ae9d4e075063 17. [reference] [navigation] wiki/Lexicon.md | bytes=2119 | sha256=d6e599c1a18ad28ed4753f6eeae7a525d87c9949a0f14a3e8f7b2a4394bd12b3 18. [reference] [navigation] wiki/Outputs-Text.md | bytes=1788 | sha256=2f71e2a6c8dbc9671ff1b4f6ab1d8084356f77655976ce04e6332d644b6ba231 19. [reference] [navigation] wiki/Outputs-UD.md | bytes=1636 | sha256=e95fa0ad10717147ef8decfd0edec3f507c44525794e17804c7cce1177489abf 20. [reference] [navigation] wiki/Positioning-Ninai-Udiron-GF-UD-No-WMF-AW-Affiliation.md | bytes=3332 | sha256=44ae0cf3d25948a06346f991b26ab1e66f2fb32e45c7f117757f4d2cc32833a0 21. [reference] [navigation] wiki/Renderer.md | bytes=1591 | sha256=ac680e013da4d0ed88ab013fc9dbb34e13db1f181db160dcc73c7fd00868b06a 22. [reference] [navigation] wiki/Repo-Map.md | bytes=3148 | sha256=912e3383db600ee0ea36fd61c02e0681d2b5f8ebe725bf650cb66c0445861379 23. [reference] [navigation] wiki/Roadmap.md | bytes=2588 | sha256=bcffae98554b68e72ba6089fd2a4f94b35546bfe7f7e75e7282bdc2b106b9bad 24. [reference] [navigation] wiki/Setup.md | bytes=2736 | sha256=524dbcfdf2691d5601a4fc1727d7ed3525b532b7d584735428a2840b6ac49323 25. [reference] [navigation] wiki/What-SemantiK-Architect-Is.md | bytes=2078 | sha256=87322206668f233d39f5dab3aadec45f4aee720b81a1ddfc56b3d299a155216b 26. [reference] [knowledge] data/raw_wikidata/README.md | bytes=1511 | sha256=e0b8e962a227a6db71b7be9fed1961b6c660d7672118c5d65d4ddb4bbb904bc5 27. [reference] [knowledge] data/reports/lexicon_coverage_report.md | bytes=591 | sha256=4d83847cc10e65a82f99b66b5d1f339d5dedc859f6ce44fc8c0d6b06c3c88aa9 28. [reference] [knowledge] docs/Technical-Reference/00-SETUP_AND_DEPLOYMENT.md | bytes=10257 | sha256=d6251a1b2846a1784636c11c9f36a32bbd80add20c39fb6ae6f8326fa85b7de5 29. [reference] [knowledge] docs/Technical-Reference/01-ENGINE_ARCHITECTURE.md | bytes=15325 | sha256=3686215728cdefa5d034a5ef462136fa67996e5125a91f09e861135d75e7e73c 30. [reference] [knowledge] docs/Technical-Reference/02-BUILD_SYSTEM.md | bytes=7508 | sha256=a58cea9ca07ff41936bde886cba6b3a6710657ae34e22c0d95c6b07cb3edb956 31. [reference] [knowledge] docs/Technical-Reference/03-LEXICON_ARCHITECTURE.md | bytes=6501 | sha256=71b52df65b3dff70ce2b6d95f1f9578ac373d149f51d405cf2c3b64956f47442 32. [reference] [knowledge] docs/Technical-Reference/04-API_REFERENCE.md | bytes=15891 | sha256=202b1b01a3d8359f8fe08add08f173ade29f2397c34c32988633d50878953b1a 33. [reference] [knowledge] docs/Technical-Reference/05-AI_SERVICES.md | bytes=7271 | sha256=4ecab865e8ab87bfe5087abadfebbac7cfaa9b5adba58b787a09c169fbb0133a 34. [reference] [knowledge] docs/Technical-Reference/06-ADDING_A_LANGUAGE.md | bytes=6347 | sha256=50cbc21623f059cefde9b27b00f204e74b3ecf8bc88da3e9cadeec6c85e8d3cf 35. [reference] [knowledge] docs/Technical-Reference/07-LINGUISTICS_REFERENCE.md | bytes=5676 | sha256=4cb3a3f652a330d99858916213da60eaebcb254c18c0e4f3c9925dbbb8de4057 36. [reference] [knowledge] docs/Technical-Reference/08-DECISION_LOG.md | bytes=5879 | sha256=1e99ced208b6d132ca57702cdd225bacdbae7f4a9163da6c57f0b77b9befc12b 37. [reference] [knowledge] docs/Technical-Reference/09-AI_CONTEXT_DUMP.md | bytes=4517 | sha256=3d8339a474ec44cce1369137180df4442e705d33ea49c845916bfef16a7f5a97 38. [reference] [knowledge] docs/Technical-Reference/10-GLOSSARY.md | bytes=3925 | sha256=ab57013ca79adad73b4555113cd157eada976ea7051d330b3b71ae57dfb50c99 39. [reference] [knowledge] docs/Technical-Reference/11-CONTRIBUTING.md | bytes=4140 | sha256=cfb9c5e11619ebfca6dd4cce2bffab6b67ccf5e3982b410ee0378a773e4f3459 40. [reference] [knowledge] docs/Technical-Reference/12-TOOLS_DASHBOARD_AND_API.md | bytes=14854 | sha256=ee22536e766d10c8b86fb2474fb373f02817bcf8e11b3615ea14f09b4063bcfd 41. [reference] [knowledge] docs/Technical-Reference/12-WIKIMEDIA_ALIGNMENT.md | bytes=5144 | sha256=b9f16d02648618a644a27ea68749d3c3962e51cf297dee759258330a43dcf1f0 42. [reference] [knowledge] docs/Technical-Reference/15-SCHEMA_ALIGNMENT_PROTOCOL.md | bytes=4648 | sha256=e51434098626f157e829cdd1e3b15df7c8e90840b03a65bcca82641590b96bd2 43. [reference] [knowledge] docs/Technical-Reference/16-DEV_TOOLS_AND_LAUNCHER.md | bytes=5433 | sha256=62e1dac74d4143f7e91f13f0225793961f3e7ce0645137e57d487be759209c88 44. [reference] [knowledge] docs/Technical-Reference/17-TOOLS_AND_TESTS_INVENTORY.md | bytes=36448 | sha256=2d2fc021339e9b90c3dbf3e7c213c347bef9477f388440975c5503a5bfcbd958 45. [reference] [knowledge] docs/Technical-Reference/A-SIMPLE-language_integration_workflow_reference.md | bytes=6962 | sha256=f0868cd30c1b8d87447c6297e4805cfaefd3b9e47908eb8525a92dc0ff67832e 46. [reference] [knowledge] docs/Technical-Reference/Abstract_Wiki_Architect_Build_and_Launch_System.md | bytes=4746 | sha256=838853ba8de6bc9ff9a34493339e3a341671965e6b0072c27d3ed4aa43ff6240 47. [reference] [knowledge] docs/Technical-Reference/ADR 005 Dual-Path Input Validation (Strict_vs_Prototype).md | bytes=4024 | sha256=17db741749771ac53c66d918a576006685cf95455e5283a2070144d76c4830e3 48. [reference] [knowledge] docs/Technical-Reference/ADR 006 Human-in-the-Loop Grammar Generation and GF Codex.md | bytes=4611 | sha256=a67e2a631589aed58eae8ebf4459881d8a4a4a12f697c618871f8009cb57b456 49. [reference] [knowledge] docs/Technical-Reference/ADR-001-GF_VERSION_STRATEGY.md | bytes=11016 | sha256=47cb5f33ea2beaa5a92629758c38a5da5f70fb1edc50fd714eaeca480ff2a775 50. [reference] [knowledge] docs/Technical-Reference/CURRENT_RUNTIME_STATUS.md | bytes=11235 | sha256=12acfc3b826a52b806764685c8cbae65c5083af0645cb8e24d4323825afd79a3 51. [reference] [knowledge] docs/Technical-Reference/DECISION_LOG.md | bytes=10573 | sha256=b70415d137d9fac9af90087c4e3119c3ca447575a79b91a674a1c423e9ea77ea 52. [reference] [knowledge] docs/Technical-Reference/ENGINE_INTERNALS.md | bytes=5763 | sha256=5a0155fe6407fc56c1f0f6051bd8a64a5a564f68ec9ceda08a82e020dd246c6a 53. [reference] [knowledge] docs/Technical-Reference/everything_matrix_orchestration.md | bytes=6890 | sha256=ec1537c6b48f048e05d0602d9c9acec2e68183d6bb2ab034797f045e9cdbcfce 54. [reference] [knowledge] docs/Technical-Reference/everythingMatrixUpgrade.md | bytes=5920 | sha256=34bccb6bc842f1e1ff50e0937a58cbbf3126a453e2359706cce3a82b8a92b532 55. [reference] [knowledge] docs/Technical-Reference/GF WordNet Structural Map.md | bytes=4120 | sha256=7cd5997caba0a642e82f6914a66261e3409b96cc51de6c72dad52d4c18b1916a 56. [reference] [knowledge] docs/Technical-Reference/GF_ARCHITECTURE.md | bytes=9027 | sha256=87c07e5535680442a5c7c0985b0f210d62209bd30ca9a02c23d365403638b897 57. [reference] [knowledge] docs/Technical-Reference/GF_Concepts.md | bytes=4206 | sha256=5c3f39f7a6d9c75076160bbc559836509593d34a289f31a34e8f155734395e69 58. [reference] [knowledge] docs/Technical-Reference/REFERENCE_LINGUISTICS.md | bytes=22112 | sha256=cf275afe945a1361d45ce90f554a1aec39ae88f23c07023fe02f86a04be7757b 59. [reference] [knowledge] docs/Technical-Reference/RESEARCH_GF_MAPPING.md | bytes=20383 | sha256=1ab4ea3c2f0b00f560ceba0708bcabe92731851f9ea1b8028e911fccd90e4e4b 60. [reference] [knowledge] docs/Technical-Reference/RGL_DISCOVERY_STRATEGY.md | bytes=5434 | sha256=1078edb280aa1a40efcfef84ebb6bcc61628a4277f95d0736c62be1c8de0862e 61. [reference] [knowledge] docs/Technical-Reference/TechnicalStatus23dec2025.md | bytes=3719 | sha256=dea99839f0ba5c37ba139fa875bf36b3557e09073aad5e1da7a446fc172addb1 62. [reference] [knowledge] docs/Technical-Reference/UPGRADE v2-0 Omni-Upgrade Architecture Specification.md | bytes=10216 | sha256=dc4a59ad519aada16e16b8d2ce728d1dd9ad3457765fdcf197357b5541159d01 63. [reference] [knowledge] docs/Technical-Reference/UPGRADE v2-0 Variable and Configuration Ledger.md | bytes=8642 | sha256=a76dcdb9011ccc287f003387b8363393397835b9274c1914e6395b6c90aa1500 64. [reference] [knowledge] docs/Technical-Reference/UPGRADE_v2.1_SPECIFICATION.md | bytes=6600 | sha256=ab574b1922e657556c99acaa6dae55d6fc0102da767d1ce42c00746e0c7dffeb 65. [reference] [knowledge] LICENSE_EXCLUSION_RATIONALE.md | bytes=2240 | sha256=5bcd71b8cf1f7e8675abbe609949ab20326befe6e9d72bbf0e3fca1b9f068ab7 66. [reference] [knowledge] README.md | bytes=9378 | sha256=dd17124d79552018a8fff1a578230e367a5f44d89e357d70fb14bbbdda26cdd5 67. [reference] [knowledge] RELEASE.md | bytes=515 | sha256=b2927c0425c3b7c20c18e8610524934ec9d34f41f01830347b26412879dc5cf2 ================================================================================================ FILE: wiki/_Footer.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 6393d670aaed7c8fe1b6a7caf54ff6b8d0e2fdb4e5a20d288b02684da33109cf CONTENT_BYTES: 267 ================================================================================================ --- **SemantiK Architect Wiki** - Navigation: use the sidebar to browse pages. - Conventions: page titles use *Title Case*; links use `[[Page Name]]` or `[[Label|File-Name]]`. - Status: if a page is draft or version-specific, it should say so at the top. --- ================================================================================================ FILE: wiki/_Sidebar.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 76b6cb73b2a55f1d0747a204525d62faab72babeb56684bce8da8d442baaef85 CONTENT_BYTES: 1035 ================================================================================================ # SemantiK Architect ## Start here - [[Home]] - [[What SemantiK Architect Is|What-SemantiK-Architect-Is]] - [[Positioning: Ninai, Udiron, GF, UD (No WMF/AW Affiliation)|Positioning-Ninai-Udiron-GF-UD-No-WMF-AW-Affiliation]] ## How it works - [[Conceptual Flow: Meaning → Text|Conceptual-Flow-Meaning-to-Text]] - [[Lexicon]] - [[Grammar]] - [[Renderer]] - [[Context]] - [[Language Coverage Strategy (Tiers)|Language-Coverage-Strategy-Tiers]] ## Quality - [[Correctness & Verifiability (Gold Standards, UD Export, Judge)|Correctness-and-Verifiability-Gold-Standards-UD-Export-Judge]] ## Using SemantiK Architect - [[Inputs: Frames|Inputs-Frames]] - [[Inputs: Ninai|Inputs-Ninai]] - [[Outputs: Text|Outputs-Text]] - [[Outputs: UD|Outputs-UD]] - [[Add a Language (High Level)|Add-a-Language-High-Level]] - [[Improve a Language (High Level)|Improve-a-Language-High-Level]] ## Developer notes - [[Setup]] - [[API Overview|API-Overview]] - [[Repo Map|Repo-Map]] ## Project - [[Decisions]] - [[Roadmap]] - [[Changelog]] - [[Glossary]] ================================================================================================ FILE: wiki/Add-a-Language-High-Level.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: ac05585c7caf28b5f2e978379caa39cf30f7aa00256ec2a8c7955b67ed174064 CONTENT_BYTES: 3826 ================================================================================================ # 15. Add a Language (High Level) ## Goal Bring a new language online with a **minimal, runnable** generation path first, then iterate toward higher quality using the project’s maturity signals. SemantiK Architect treats “language support” as a **measured state** (not a checkbox): the system computes per-language maturity and uses it to decide what is runnable and which strategy to use (high-quality vs safe-mode). --- ## Step 1 — Choose an initial quality path (Tier choice) Pick the best available strategy for the language: - **Tier 1 (High quality):** choose this when the language already has strong, mature grammar support (e.g., a robust GF/RGL-class path). - **Tier 3 (Guaranteed availability / Safe Mode):** choose this for under-resourced languages so the language is available quickly via the **weighted topology / factory** approach, then improve later. - **Tier 2 (Manual overrides):** use this as the bridge when you want to improve phrasing/behavior for a specific language without waiting for full Tier 1 coverage. --- ## Step 2 — Create the language “home” (language code and identity) Use a **two-letter ISO-2 code** as the **internal language key** (folder names, core data identity, and internal registry expectations). Important nuance: - **Internal identity** is ISO-2 (the consistent key the system reasons about). - **External/public identifiers** (if you expose them) may require explicit mapping and do not have to be “ISO-2 everywhere.” --- ## Step 3 — Seed the minimum Lexicon so the language is runnable A language becomes runnable only when it clears a minimum **lexicon maturity** threshold. Minimal expectations (example policy): - **Minimal:** `core.json` exists and contains enough entries to generate basic sentences. - **Functional:** adds essential domain shards (often starting with `people` for biographies). Bootstrapping workflow (high level): 1. Create the language lexicon directory. 2. Seed `core.json` (manually or with helper tooling). 3. Run an audit/index step so the system discovers the new language. Guardrail: - If the “core seed” is too small, the system should treat the language as **non-runnable** to avoid brittle generation paths. Optional: - A “lexicon bootstrapper”/helper can generate a starter `core.json` specifically to prevent “empty dictionary” languages. --- ## Step 4 — Register + audit via the Everything Matrix (no hardcoded lists) Do **not** hardcode language lists. The rule is: - **Create the expected file structure**, then - Let the indexing/audit step pick up the language and compute maturity. The Matrix acts as the “central nervous system” for autonomous decisions (e.g., whether to run Tier 1 or fall back to safe-mode). --- ## Step 5 — Add at least one correctness checkpoint (so it can’t silently regress) “Done” is not “it compiles.” Add at least one baseline check that can fail if quality regresses. Minimum recommendation: - Add **one gold-standard example** for the language and ensure the automated evaluator (“Judge”) runs it. Project policies sometimes use a similarity threshold (e.g., **0.8**) to block regressions; treat the exact number as a project-level policy that can evolve. If your language work introduces new factory-path constructions, keep the **validation/export mapping** (e.g., UD view) aligned so the output remains auditable. --- ## Definition of “done” (for a first PR) A new language is “added” when: - It exists under the **internal ISO-2 key** (with any external mapping handled explicitly if needed) - It has at least a **minimal lexicon seed** that makes it runnable - It is discovered by the **Everything Matrix** audit/index step - It has at least **one gold-standard case** so quality can be tracked over time ================================================================================================ FILE: wiki/API-Overview.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: e161a286f53d2a1a8a2aea8c462d3bb4554a3ee853b61c375a39a08a4e375286 CONTENT_BYTES: 2275 ================================================================================================ # API Overview SemantiK Architect is accessed through a **versioned HTTP API**. The **stable convention** is to keep client-facing endpoints under the **`/api/v1`** prefix. > Note: some deployments also mount the UI/API under a **base path**. Treat the base path as a deployment choice, not an API rule. ## Base path (deployment) In packaged deployments, the UI may be served under a base path and the API is served under: - **`{BASE_PATH}/api/v1`** Example (legacy): UI under `/abstract_wiki_architect`, API under `/abstract_wiki_architect/api/v1`. ## Public vs admin endpoints (auth boundary) - **Public read endpoints** may be accessible without admin credentials (deployment-configurable). - **Admin endpoints** require authentication (API key/token). - The **Tools API** is **admin-only**. ## Core generation endpoint (meaning → text) ### `POST /api/v1/generate/{lang}` - `lang` is passed as a **path parameter**. **Two supported input shapes (conceptually):** - **Strict / production path:** a flat “frame” JSON (example: `BioFrame`). - **Prototype / experimental path:** a recursive Ninai-style tree (`UniversalNode`). **Response shape (stable envelope):** - `surface_text`: the generated text - `meta`: lightweight provenance (e.g., engine/adapter/strategy) ## Session/context support (multi-sentence coherence) Requests may include a session identifier header to enable discourse behavior across sentences (example: `X-Session-ID`). ## Discovery endpoints (used by the UI) Some deployments/UI flows rely on discovery endpoints such as: - `GET /api/v1/languages` returning a list of languages (for a language selector). - `GET /api/v1/entities/...` returning a defined entity schema (for browsing/search). If your current build does not expose these yet, treat them as the **target contract** for UI/API alignment. ## Tools API (GUI-driven operations; admin-only) Tools are executed via an **allowlisted registry** (no arbitrary command execution). - `GET /api/v1/tools/registry`: lists tool metadata - `POST /api/v1/tools/run`: runs a tool by `tool_id` and returns a stable execution envelope (trace/output/events/exit code) **Security note:** do not pass secrets in tool args; args may be echoed in responses and/or surfaced in UI logs. ================================================================================================ FILE: wiki/Changelog.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: faf9abec1521092ff6cfdb5cec7bc250142067733c1c285e63fa84bd3eb38245 CONTENT_BYTES: 1973 ================================================================================================ ## 22. Changelog ### Unreleased (SemantiK Architect) * Project renamed from **“Abstract Wiki Architect” → “SemantiK Architect”** (independent project; no WMF/Abstract Wiki affiliation). * Wiki scope reduced to a simpler v1 (high-level pages first; deeper material postponed). ### v2.5 (Docs baseline: setup/deploy + operator workflow) * Documented the **Windows + WSL2 hybrid** dev model (Windows for editing/frontend, Linux/WSL for backend + GF), driven by the GF `libpgf` Linux dependency. * Standardized prerequisites (WSL2, Ubuntu, Docker Desktop, VS Code WSL extension, Node 18+). * Clarified **deployment via docker-compose**, including base-path conventions for UI and versioned API prefix. * Formalized config/ops knobs in `.env`, including explicit **AI tool gating**. ### v2.1 (Architecture and build system clarified/expanded) * Consolidated the system model as a **4-layer architecture**: Lexicon, Grammar, Renderer, Context. * Defined the **dual-path input** (Strict Frames vs Prototype Ninai/UniversalNode) and dual output (Text or UD/CoNLL-U). * Codified the **3-tier language strategy** (Tier 1 “High Road”, Tier 2 overrides, Tier 3 weighted-topology factory) for long-tail coverage. * Made the build “self-aware” via the **Everything Matrix** (filesystem scanning → maturity scoring → build strategy). ### v2.0 (Omni-upgrade milestone: expansion beyond sentence-level generation) * Stated the core shift: from a **sentence-level rule-based engine** to a **context-aware, interoperable, AI-augmented platform**. * Introduced the “7 pillars” roadmap: Ninai bridge, UD exporter, discourse planner, automation agent, interactive QA, weighted topology factory, learned micro-planning. * Introduced a unified **“Check, Build, Serve”** pipeline and a single operational entry point for developers (`manage.py`). * Established the **Everything Matrix** as a dynamic registry replacing static language lists/config. ================================================================================================ FILE: wiki/Conceptual-Flow-Meaning-to-Text.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: b584eb543da5bcb0216ea520f937fd0cba9410d384abaa78d42ec6d7b31310a9 CONTENT_BYTES: 2456 ================================================================================================ # 4. Conceptual Flow: Meaning → Text SemantiK Architect is **generation-first**: it does not “interpret” text. It takes an **unambiguous meaning payload** and renders it into surface language. ## Flow (at a glance) **Meaning (Frame or Ninai)** → **Normalize meaning** → **Apply discourse context (optional)** → **Pick language strategy (tier)** → **Realize text (lexicon + grammar)** → **Export (text, optionally UD)** → **Quality checks (optional)** ## Step-by-step 1. **Provide meaning (input)** SemantiK Architect accepts meaning in two shapes: - a **strict, flat semantic frame** (stable production input) - a **recursive Ninai-style object tree** (more expressive, experimental input) 2. **Normalize meaning (adapter stage)** Inputs are normalized into an internal “intent” representation: - validate required fields / structure - resolve defaults and normalize naming - if the input is Ninai-style, an adapter walks the object tree and converts it into the internal intent/frame representation (a recursive object-walker approach, not text parsing) 3. **Use context for multi-sentence coherence (optional but important for naturalness)** If a session is active, the system can track what entity is “in focus” and apply simple discourse decisions (e.g., reducing repeated names via pronouns when appropriate). 4. **Select a language strategy (coverage vs precision)** To cover both high-resource and long-tail languages, the renderer selects a tiered strategy: - a higher-precision **rule-based** path (when strong grammar resources exist) - a broader-coverage **factory/topology-based** fallback (when they don’t) - optional manual overrides can take precedence when available 5. **Realize the sentence (meaning → surface text)** Rendering is the assembly step that combines: - **lexical choice** (words and their properties) - **grammar / linearization** (how those words become a correct sentence in the target language) 6. **Export (outputs)** The primary output is **natural language text**. Optionally, the system can also output a **Universal Dependencies (CoNLL-U) view** for validation and evaluation. 7. **Close the loop with quality checks (optional workflow)** A QA loop can compare generated output against a **gold standard** and flag regressions (including automated reporting workflows), so improvements remain stable over time. ================================================================================================ FILE: wiki/Context.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: b43ca7a23af89688ad77573eac023612be4546a836dc1b92862d1f627fd4b193 CONTENT_BYTES: 2462 ================================================================================================ # Context ## What “Context” means in SemantiK Architect Context is the **memory layer** that lets SemantiK Architect move beyond generating isolated sentences and instead produce **coherent multi-sentence text**. In the architecture, it is explicitly the layer responsible for **session state** used for **discourse planning**. ## Why it exists Without context, the system repeats full names and produces text that feels unnatural. The docs illustrate this with a two-sentence example where repeating “Marie Curie” is correct but stylistically poor. Context exists to address that by tracking which entity is currently “in focus” so the system can choose more natural referring expressions. ## What it tracks (conceptually) SemantiK Architect keeps a lightweight “discourse state” that includes: * **Mentioned entities** * **Which entity is salient / in focus** * **Last mention and topic selection signals** This is intentionally “light but real”: enough structure to handle multi-sentence needs without adopting a heavy formal discourse/semantics framework. ## What it enables (user-visible effects) ### 1) Pronominalization (avoiding repetition) When the subject of the current sentence matches the current “in-focus” entity from the session, the system can replace a repeated name with a pronoun (“Swap Name → Pronoun”). ### 2) Basic coherence across sentences Context supports “discourse planning” decisions such as: * pronouns vs full names, * topic markers vs default word order, * basic ordering of information. ## What Context is *not* * It is **not** a full semantic/discourse formalism (the docs explicitly reject heavy approaches for the initial implementation). * It is **not** “style generation” by itself; it’s a control layer that helps the renderer make consistent choices across sentences. ## How it fits with the rest of the system * The **Renderer** decides what sentence to produce now. * **Context** provides the “what has been said / what is in focus” memory so the renderer can produce a better next sentence. * The “Summary of Systems” explicitly pairs the discourse planner with **Centering Theory / coreference** as the conceptual grounding. ## Where this is going (safe, high-level roadmap) The docs describe future upgrades toward more powerful discourse models for anaphora and topic shifts, enabling longer multi-sentence texts while keeping coherence. ================================================================================================ FILE: wiki/Correctness-and-Verifiability-Gold-Standards-UD-Export-Judge.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 17c7a2601aaed1482f76128bd9797c84655dfd2d96633778fb6b3235a38c07bd CONTENT_BYTES: 2639 ================================================================================================ # 10. Correctness & Verifiability (Gold Standards, UD Export, Judge) ## Why this matters SemantiK Architect is designed to scale to **hundreds of languages**, which makes “manual checking” impossible as a primary quality strategy. The system therefore treats correctness as something that must be **measurable, repeatable, and enforceable**. ## What “verifiable” means here “Verifiable” means SemantiK Architect doesn’t just generate text—it also produces **checkable evidence** that the output remains consistent and linguistically sound across time and across languages. One key source of high-confidence correctness is the rule-based grammar path (GF/RGL) for high-resource languages, described as providing “verifiable correctness.” ## The three mechanisms ### 1) UD Export (standards-based validation surface) SemantiK Architect can export a **Universal Dependencies (UD)** representation (CoNLL-U mapping) so outputs can be validated and compared in a standardized way. The documentation sets a strict rule: every syntactic constructor must have a corresponding UD tag mapping—so validation coverage can’t silently drift. ### 2) Gold Standards (ground truth expectations) A “Gold Standard” is a curated set of **verified intent → expected text** pairs used as a regression baseline. Major changes are expected to be validated against this dataset. The spec also describes ingesting a test suite as ground truth (migrated from Udiron in the original lineage), reinforcing that Gold Standards are meant to be **externalized, reusable, and stable**. ### 3) The Judge (automated regression gate) “The Judge” is the evaluation loop that compares current outputs to Gold Standards and flags regressions. The docs describe a hard gating concept: PRs can be blocked if the Judge score drops below a threshold (0.8 is explicitly mentioned). The same quality loop is also described as closing the loop operationally by validating and (optionally) auto-reporting failures. ## How to interpret failures (high level) When a regression is detected, it usually points to one of four buckets: * **Lexicon gap** (missing/incorrect word data), * **Grammar/realization issue** (wrong structure or morphology), * **Mapping/validation gap** (UD mapping missing or inconsistent), * **Renderer/context behavior change** (surface text changed unintentionally). The purpose of this page in the wiki is to make one idea clear: **SemantiK Architect treats correctness as a first-class product feature, enforced by standards export + gold truth + automated judging**, not as a best-effort manual process. ================================================================================================ FILE: wiki/Decisions.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 9cc35f8c8294ca251148436b3f56a4438f91580c047a4d0594c0acdd82f38f12 CONTENT_BYTES: 2718 ================================================================================================ # 20. Decisions ## Purpose This page records the **product-defining choices** that explain *why SemantiK Architect works the way it does*, and what must stay stable when the system evolves. (Notes: the source docs use the former name “Abstract Wiki Architect”; the decisions below remain applicable under the SemantiK Architect rename.) --- ## Canonical decisions (high level) ### 1) Data-driven registry instead of hardcoded configuration (“Everything Matrix”) * **Problem:** hardcoded language lists drift from reality. * **Decision:** a dynamic registry is rebuilt by scanning what exists on disk before each build. * **Why:** it becomes the single source of truth; adding a language becomes “add the files.” ### 2) Hybrid language strategy instead of a single approach * **Decision:** a tiered strategy: high-resource languages use GF/RGL; under-resourced languages use a weighted-topology approach (from the Udiron lineage) for coverage. * **Why:** it explicitly manages the quality vs coverage tradeoff rather than pretending one method fits all. ### 3) Standards interoperability as a first-class requirement * **Decision:** design for interoperability with **Ninai** (meaning protocol) and **UD** (validation/export surface), framed as a core architecture choice. ### 4) Discourse/context is a core capability (not an afterthought) * **Decision:** maintain session context (e.g., for pronouns/discourse planning) as a key architectural pillar. ### 5) Quality is enforced via Gold Standards + an automated Judge * **Decision:** regressions are checked against Gold Standard data (explicitly ingesting the Udiron test suite lineage) and gated by a similarity threshold (0.8 is called out). ### 6) AI is used as a scale multiplier for generation and evaluation (as designed in v2.0) * **Decision:** specialized agents (“Architect” for generating Tier-3 grammar assets and “Judge” for validation + issue filing). * **Note for your wiki:** if your current SemantiK direction reduces/removes build-time AI, keep this decision as “historical” or “optional mode,” not as the default. ### 7) “No arbitrary execution” in tooling * **Decision:** operational tools are run through a strict allowlist registry rather than arbitrary commands, to keep the system safe and predictable. --- ## How new decisions should be recorded When changing any of the above (e.g., removing AI from builds, changing tier strategy, changing gating thresholds), add an entry that states: * Context / problem * Decision * Why * Consequences (what becomes easier/harder) This keeps SemantiK Architect understandable without turning the wiki into low-level technical documentation. ================================================================================================ FILE: wiki/Glossary.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 484745566f66e08a76ae0d4060f0e02a367b10996863b294aa7927fe88307403 CONTENT_BYTES: 5599 ================================================================================================ # Glossary (SemantiK Architect) > **Naming note:** *SemantiK Architect* is the current name of the project (formerly **Abstract Wiki Architect** in earlier docs and some filenames). > > **Independence note:** SemantiK Architect is **not affiliated with** or developed in collaboration with **WMF / Abstract Wikipedia / Abstract Wiki**. --- ## Project + related tools (boundaries) - **SemantiK Architect**: A multilingual **renderer** that turns structured meaning into readable natural-language text, designed to scale across high-resource and long-tail languages. - **Abstract Wiki Architect (legacy name)**: Former project name that may still appear in v2 docs, historical notes, and some repo artifacts. - **GF (Grammatical Framework)**: A grammar formalism/toolchain used (where available) to define and run high-quality grammars. - **RGL (Resource Grammar Library)**: GF’s standard library for morphology/syntax in a set of languages; often the basis of the “high-quality” path when applicable. - **UD (Universal Dependencies)**: A dependency-grammar standard used here mainly as a **validation/evaluation view** (often via CoNLL-U export). - **Ninai (meaning representation)**: A recursive, constructor-tree style meaning representation that SemantiK Architect can accept (typically via an adapter/bridge). - **Udiron / weighted-topology approach**: A coverage-oriented strategy that uses configurable ordering/topology rules to scale to many languages when full expert grammars are not available. --- ## Meaning representations (inputs) - **Semantic Frame (Frame)**: A language-agnostic, usually flat JSON object expressing intent (e.g., “bio”, “event”) meant to be stable and easy to validate. - **BioFrame / EventFrame**: Concrete frame “shapes” (types) used by the strict/production input path. - **Recursive meaning tree**: A nested, tree-shaped input form (often Ninai-style) used for richer meaning composition and prototyping. - **Adapter / Bridge**: A component that converts an external meaning format (e.g., Ninai-style trees) into the renderer’s internal representation. --- ## Core generation concepts - **Lexicon**: Vocabulary entries (words/lemmas + properties) used during realization. - **Grammar**: Rules that handle morphology (inflection) and syntax/word order. - **Renderer**: The “assembly” layer that maps meaning to a sentence plan and realizes it into output text (and optional validation views). - **Context**: Discourse/session state used to keep multi-sentence generation coherent (e.g., reference, repetition, pronouns). Implementation details may vary. --- ## GF / linguistics terms (only when relevant) - **Abstract syntax**: Language-independent structures defining *what* can be expressed. - **Concrete syntax**: Language-specific rules defining *how* to express it in a given language. - **Linearization**: Turning an abstract structure/tree into a final surface string. - **Morphology**: Word-level inflection (plural, tense, case, agreement, etc.). --- ## Language coverage strategy - **Tier 1 (High quality)**: Uses strong, mature grammar resources (often GF/RGL-class) for best grammatical correctness. - **Tier 2 (Manual overrides)**: Targeted language-specific improvements/overrides layered on top of other tiers. - **Tier 3 (Factory / Safe Mode)**: Coverage-first grammars driven by configurable topology/ordering so a language is available quickly, even if nuance is limited. - **Weighted topology**: A Tier 3 mechanism that uses configurable weights/rules to decide ordering (e.g., SVO/SOV tendencies) rather than hardcoded templates. --- ## Build/runtime and configuration terms - **Everything Matrix**: A computed registry of what exists per language (assets, maturity signals, runnable status) used to avoid hardcoded language lists. - **Two-phase build (concept)**: A build strategy that verifies components first and then links/aggregates them, to avoid partial overwrites and “last one wins” behavior. - **PGF (Portable Grammar Format)**: The compiled GF artifact loaded at runtime. Filenames may still include legacy naming in some repos. --- ## Outputs + evaluation - **Text output**: The default natural-language surface string. - **CoNLL-U output**: Optional export of a UD-style dependency representation for validation/evaluation. - **Gold Standard**: A curated set of reference examples used to track quality over time. - **Judge**: An automated evaluator that compares output against Gold Standard data and flags regressions (optionally integrated with CI and issue tracking, depending on setup). --- ## Data organization - **Domain sharding**: Splitting vocabulary into topic-focused files (e.g., `core`, `people`, `science`) so systems can load what they need. - **Wikidata QID mapping**: Grounding lexicon entries to stable identifiers when available, to keep meaning aligned across languages. --- ## Software architecture terms (lightweight) - **Ports & adapters (hexagonal architecture)**: A structuring approach that separates core generation logic from infrastructure (APIs, storage, tooling). --- ## Automation / agent terms (status may vary) - **Automation agent (e.g., “Architect”, “Surgeon”)**: Names used for optional tooling intended to assist with authoring/repairing data or grammars. These are **not required** for the core deterministic generation path, and whether they are active depends on the project’s current workflow. - **Frozen system prompt**: A fixed prompt template concept used to keep automated outputs consistent when agents are used. ================================================================================================ FILE: wiki/Grammar.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: d06b24c10f0943488dee118bc721e6f7e05112ca03602208b78a205be595aaa4 CONTENT_BYTES: 3167 ================================================================================================ # Grammar The **Grammar** layer is the rule system that turns structured meaning into **well-formed text** in a target language. In SemantiK Architect, “grammar” covers two things: - **Morphology** (inflection): how words change form (agreement, conjugation, declension). - **Syntax** (structure + word order): how parts combine into phrases/clauses and how they are ordered. This is the “rules” layer of the overall architecture: it defines morphology and syntax/word order, and it supports a **hybrid** approach that balances quality and coverage. --- ## Why this layer exists SemantiK Architect aims for: - **Determinism** (same input → same output), - **Broad language coverage** (including the long tail), - **Reasonable grammatical quality** even when a language lacks a full expert grammar. The grammar layer is where those tradeoffs are managed: high precision when strong resources exist, and graceful degradation when they don’t. --- ## The hybrid strategy (tiers) SemantiK Architect uses a tiered approach to grammar resources: - **Tier 1 (High Road): expert grammars** For high-resource languages, use expert-grade grammar resources (GF/RGL-class) for richer morphology and more precise sentence realization. - **Tier 2: manual / curated overrides** When a community or project-maintained grammar exists, it can override both Tier 1 and Tier 3 for that language to improve naturalness and correctness. - **Tier 3 (Factory): weighted-topology fallback** For under-resourced languages, use a configuration-driven fallback based on **weighted topology / dependency-role ordering**. This exists to avoid “missing language” dead ends and ensure the system can still produce a grammatical-enough sentence. See also: [[Language Coverage Strategy (Tiers)|Language-Coverage-Strategy-Tiers]] --- ## Grammar matrices and language cards To scale efficiently, grammar knowledge is organized at two levels: - **Family grammar matrices** Reusable defaults shared across related languages (e.g., Romance, Slavic). They capture broad paradigm space and common patterns. - **Language cards** Language-specific overrides for quirks that shouldn’t live in the shared family matrix (e.g., special elision or clitic behavior). This keeps the system modular: broad coverage from shared structure, with targeted exceptions where needed. --- ## Relationship to other layers - **Lexicon → Grammar** The lexicon supplies lemmas and features (e.g., gender/number) that the grammar needs to inflect and agree. See: [[Lexicon]] - **Grammar → Renderer** The renderer selects a grammar strategy (tier) and uses it to realize a sentence plan into text. See: [[Renderer]] - **Context ↔ Grammar** Context can change *what* gets expressed (e.g., pronoun vs name); grammar ensures the chosen form fits correctly in the sentence. See: [[Context]] --- ## What this page is (and is not) This page explains what “Grammar” means in SemantiK Architect and how it supports coverage and quality. It is **not** a build guide or an API reference. - Setup/build: [[Setup]] - API surface: [[API Overview|API-Overview]] ================================================================================================ FILE: wiki/Home.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 2d8f622c87e1d03b67492b5505007dde746ba652ed18610ebee763157403e4fe CONTENT_BYTES: 3567 ================================================================================================ # SemantiK Architect SemantiK Architect is an **independent multilingual text renderer**: it turns **structured meaning** into **encyclopedic natural language**, with a focus on scaling to the **long tail** of languages (including under-resourced ones). > Note: SemantiK Architect is **not affiliated with WMF / Abstract Wikipedia** and is **not developed in collaboration** with the Abstract Wiki team. It is the continuation/rename of “Abstract Wiki Architect” as an independent project. --- ## What it does SemantiK Architect: - **Accepts meaning** in a structured form (e.g., *Frames* or *Ninai-style* recursive objects). - **Generates text** in a target language. - Optionally **exports a linguistic view** (e.g., Universal Dependencies / CoNLL-U) to support validation and evaluation. --- ## How it works (conceptual) At a high level, the system is organized into four conceptual parts: - **Lexicon** — vocabulary (words + properties), grounded to stable identifiers when possible. - **Grammar** — rules for inflection and word order. - **Renderer** — turns the input meaning into a sentence plan and realizes it as text. - **Context** — tracks discourse state across sentences (e.g., reference, focus, pronouns). To scale language coverage, SemantiK Architect uses a **tiered strategy**: - **Tier 1**: highest-quality grammars where available (e.g., GF/RGL-class resources). - **Tier 2**: curated/manual improvements and overrides. - **Tier 3**: automated “factory” grammars (e.g., topology/word-order driven), inspired by UD/Udiron-style approaches, to avoid missing-language dead ends. --- ## Where it stands vs related tools - **Ninai**: a meaning representation / protocol. SemantiK Architect is a **renderer** that can consume it. - **GF**: a grammar technology. SemantiK Architect **uses/hosts** grammars as part of a larger end-to-end renderer. - **Udiron-style topology**: a strategy for rapid scaling via word-order/topology configuration; used as a **coverage path**. - **UD (Universal Dependencies)**: a cross-linguistic interface useful for **evaluation and verification** (not a replacement for generation). See: [[Positioning: Ninai, Udiron, GF, UD (No WMF/AW Affiliation)|Positioning-Ninai-Udiron-GF-UD-No-WMF-AW-Affiliation]] --- ## How to read this wiki Recommended order: 1. [[What SemantiK Architect Is|What-SemantiK-Architect-Is]] 2. [[Positioning: Ninai, Udiron, GF, UD (No WMF/AW Affiliation)|Positioning-Ninai-Udiron-GF-UD-No-WMF-AW-Affiliation]] 3. [[Conceptual Flow: Meaning → Text|Conceptual-Flow-Meaning-to-Text]] 4. Components: [[Lexicon]] → [[Grammar]] → [[Renderer]] → [[Context]] 5. [[Language Coverage Strategy (Tiers)|Language-Coverage-Strategy-Tiers]] 6. [[Correctness & Verifiability (Gold Standards, UD Export, Judge)|Correctness-and-Verifiability-Gold-Standards-UD-Export-Judge]] 7. Using pages: [[Inputs: Frames|Inputs-Frames]] / [[Inputs: Ninai|Inputs-Ninai]] / [[Outputs: Text|Outputs-Text]] / [[Outputs: UD|Outputs-UD]] Project pages: - [[Roadmap]], [[Changelog]], [[Decisions]], [[Glossary]] --- ## What this wiki is (and is not) - **This wiki is** a high-level explanation of what SemantiK Architect is, how it fits in its ecosystem, and how the main concepts connect. - **This wiki is not** a full technical manual (setup scripts, build internals, exhaustive API reference). Those belong in minimal “Developer Notes” pages and/or repository docs. If you’re here to implement or integrate quickly, start at: [[API Overview|API-Overview]] and [[Repo Map|Repo-Map]]. ================================================================================================ FILE: wiki/Improve-a-Language-High-Level.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 0647d1388d6e6c3aeac503a21c7fb2b3258abc622ffe2f26dbe0dd6ce8f3bdc7 CONTENT_BYTES: 5474 ================================================================================================ # Improve a Language (High Level) SemantiK Architect can generate usable sentences for many languages, but languages are not all supported at the same quality level. “Improving a language” means moving it upward on a quality ladder while keeping output stable, predictable, and buildable. This page explains the **levers you can pull** (vocabulary, grammar, and QA), and the **typical path** from “it works” to “it feels native”. --- ## 1) The 3 support tiers (how a language is handled) SemantiK Architect uses a three-tier approach: - **Tier 1 — High Road (RGL-quality)** - Best grammatical quality and richest morphology. - Preferred when the language has strong, mature grammar coverage. - **Tier 2 — Manual Overrides** - Community-contributed or project-contributed grammar improvements. - Takes precedence when present (it can “override” the other tiers). - **Tier 3 — Safe Mode (Factory)** - A fallback designed to avoid “we don’t support that language”. - Prioritizes *always returning a sentence*, even if style/nuance is simpler. --- ## 2) What makes a language “better” in practice Language quality is usually felt through three things: ### A. Vocabulary coverage (Lexicon) A language improves quickly when it has the right words available *in the right places*. SemantiK Architect organizes vocabulary in **domain shards** (not one giant dictionary). Typical shards include: - **core**: the skeleton words you need for almost any sentence (copulas, pronouns, articles, connectors) - **people**: professions, roles, relations (so biographies sound correct) - **geography**: countries, demonyms, adjectives - **science**: specialized terminology ### B. Grammar behavior (how meaning becomes a sentence) Even with good words, a language needs reliable sentence-building rules: - Word order that fits the language’s typology - Agreement and inflection that feels consistent - Fewer “robotic” constructions as quality rises Tier 2 contributions (manual overrides) are the typical bridge to make a language feel more natural before it ever becomes Tier 1. ### C. Quality assurance (staying good over time) A language is “improved” only if it stays improved. QA is how you prevent regressions: - Reference examples (“gold standard” sentences) - Regression checks when changing lexicon/grammar - Clear signals when output quality drops --- ## 3) How SemantiK decides what to do with a language SemantiK Architect relies on a central inventory (“the brain”) that: - Discovers what language assets exist - Scores maturity/readiness - Decides whether the language should run Tier 1, Tier 2, or Tier 3 - Can downgrade to Safe Mode when a language is too incomplete (to avoid brittle builds) The key idea: the system is **data-driven**, not a manually curated list of languages. --- ## 4) The improvement loop (recommended workflow) ### Step 1 — Make sure the “skeleton” exists Start with the minimum vocabulary needed to form basic clauses (core shard). If the language can’t reliably express “X is Y”, everything else will feel broken. ### Step 2 — Make biographies work end-to-end Biographies are often the first “real” use case. Add/expand: - professions, roles, relations (people shard) - nationality and demonyms (geography shard) ### Step 3 — Fix missing words as they surface When generation fails because a word is missing, treat it as a signal: - Add the missing entry in the appropriate shard - Keep shards small and meaningful rather than dumping everything into one file ### Step 4 — Improve grammar feel with Tier 2 (manual overrides) When output is grammatical but awkward: - Add targeted manual grammar improvements (Tier 2) - Focus on the constructions that appear most often (biographies, basic relations, simple events) ### Step 5 — Graduate to Tier 1 where possible Some languages may eventually rely primarily on Tier 1-quality grammar coverage. This is a longer path, but it’s the highest ceiling. ### Step 6 — Protect the gains with QA Once a language looks good: - Add representative examples to your reference set (“gold standard”) - Use automated checking so “it used to work” doesn’t become a recurring problem --- ## 5) What you can contribute (non-technical categories) - **Lexicon contributions** - Add missing words in the right shard - Expand coverage in domains that matter (people/geography first) - **Grammar contributions** - Improve the most-visible sentence patterns first - Provide Tier 2 “overrides” that reduce awkward phrasing - **QA contributions** - Add a small set of high-signal reference sentences - Track regressions and decide what “good enough” means for each tier --- ## 6) Practical “definition of done” (for a language milestone) A language can be considered “meaningfully improved” when: - It reliably generates key sentence types (especially biographies) without missing-word failures - It has a functional lexicon skeleton plus the main domain shards needed for your targets - It has at least a minimal QA baseline (so improvements don’t evaporate) --- ## 7) Optional: measuring progress (simple mental model) Think of progress as: 1) **Coverage** (can we say it?) 2) **Naturalness** (does it sound right?) 3) **Durability** (does it stay right after changes?) Tier 3 gets you coverage fast, Tier 2 improves naturalness, and QA makes it durable. ================================================================================================ FILE: wiki/Inputs-Frames.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 776f70ee6a76ab8edf7bd2ba26631addf2c7974f761e989d08da1b41eac6b9b4 CONTENT_BYTES: 1783 ================================================================================================ # Inputs: Frames Frames are the **strict**, **flat JSON** input format used by SemantiK Architect for stable generation. In the system’s dual-path design, Frames correspond to the **“Strict Path”** (validated internal frame objects such as `BioFrame`) intended for production reliability. --- ## What a “Frame” is A Frame is a compact, structured way to express **meaning/intention** without tying it to any particular language’s surface grammar. Think of it as: **“what you want to say”**, expressed with a small set of named fields that the renderer can reliably turn into text. --- ## How SemantiK Architect interprets Frames - A Frame is identified by a **`frame_type`** field (e.g., `"bio"`). - The remaining fields are the **arguments/slots** needed to express that meaning (e.g., `name`, `profession`, etc.). - The request is routed to the corresponding strict handler and **validated** before generation. --- ## Example: BioFrame (illustrative) This example shows the “shape” of a typical Frame request: ```json { "frame_type": "bio", "name": "Alan Turing", "profession": "computer scientist", "nationality": "british", "gender": "m" } ```` (Other frame types will define different required/optional slots.) --- ## When to use Frames vs Ninai Use **Frames** when you want: * a **simple, stable contract** for upstream systems you control, * **strict validation** and predictable behavior, * an intentionally “flat” meaning form that is easy to author and debug. Use **Ninai** when you want: * a **recursive object-tree** meaning representation (more expressive), * a format that can be **adapted into internal frames** through the Ninai bridge. See also: * [[Inputs: Ninai|Inputs-Ninai]] * [[API Overview|API-Overview]] ================================================================================================ FILE: wiki/Inputs-Ninai.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: bfba15dabfe14fc6aad44a8c20af8842808b4067249e73ea8672588b4c146b29 CONTENT_BYTES: 1754 ================================================================================================ ## 12. Inputs: Ninai **Ninai** is a **meaning representation**: a structured way to express “what you want to say” before choosing any particular language. In the docs, it’s described as **recursive JSON object trees** (constructor-style structures), not plain text. ### Why SemantiK Architect supports Ninai SemantiK Architect uses Ninai as its **interoperability input**: a way to accept rich, language-independent meaning in a form that can be rendered across many languages. This is explicitly part of the v2 pillars (“Ninai Bridge”). ### How it fits in the product (high level) * Ninai is the **prototype/experimental input port** into the Renderer (the “UniversalNode / recursive Ninai JSON” path). * SemantiK Architect then **maps** that Ninai tree into its internal intent structures (e.g., Bio/Event frames) so it can generate text deterministically. * The key conceptual point: Ninai is the **meaning layer**, and Architect is the **renderer** that turns it into text. ### What Ninai looks like (conceptually) * A **tree** where each node states a “constructor/function” plus its arguments. * Because it’s recursive, the system treats it as a **meaning tree** and walks it (not a string to parse). ### When to use Ninai vs Frames * Use **Ninai** when you want: portability, richer semantics, experimentation, and a single meaning structure that can be rendered in many languages. * Use **Frames** when you want: a stable, constrained “production” input shape (the docs explicitly separate strict vs prototype paths). *(Project note: SemantiK Architect is independent and not affiliated with WMF/Abstract Wiki; Ninai is treated here as an external meaning format that SemantiK Architect can consume.)* ================================================================================================ FILE: wiki/Language-Coverage-Strategy-Tiers.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 03ef5d2f259614646500810b5b347d496391d5f7e5dae8d64b55ae9d4e075063 CONTENT_BYTES: 1723 ================================================================================================ # 9. Language Coverage Strategy (Tiers) SemantiK Architect uses a **three-tier strategy** so it can deliver **high quality where possible** and still provide **broad language coverage** (the “long tail”) without blocking on writing perfect grammars for every language. ## Why tiers exist * **Tier 1** maximizes linguistic quality (best output). * **Tier 3** maximizes coverage and robustness (never “no language support”). * **Tier 2** is the human/community override layer that can improve output without waiting for upstream libraries. ## The tiers ### Tier 1 — “High Road” (Best quality) Uses the GF Resource Grammar Library for languages where it’s strong (high-resource languages). ### Tier 2 — Manual contributions (Overrides) Community or project-maintained grammars that aren’t in the official RGL yet, but are better than automated stubs. If present, this tier **overrides** both Tier 1 and Tier 3. ### Tier 3 — “Weighted Factory” (Coverage + safety) An automated fallback that uses **weighted topology sorting** (adapted from Udiron) so word order can be configured rather than hardcoded templates. The intent is that the system remains usable across many word-order types and **doesn’t fail closed**. ## How a language’s tier is chosen (at a high level) A central registry (“Everything Matrix”) audits what exists on disk (grammar readiness, lexicon readiness), assigns a score, and selects a build strategy: * **High score → Tier 1** * **Low score → Tier 3** (graceful degradation / “safe mode”) This is also the meaning of “Hybrid Factory”: combining expert grammars (Tier 1) with automated simplified grammars (Tier 3) to reach full coverage. ================================================================================================ FILE: wiki/Lexicon.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: d6e599c1a18ad28ed4753f6eeae7a525d87c9949a0f14a3e8f7b2a4394bd12b3 CONTENT_BYTES: 2119 ================================================================================================ # 5. Lexicon ## What the Lexicon is The Lexicon is SemantiK Architect’s **vocabulary layer**: the words (and key linguistic properties) the system needs to express meaning in a specific language. It is designed to support **300+ languages** without becoming a single unmaintainable dictionary. ## Core principles * **Grounded meaning:** entries are traced back to **Wikidata QIDs** so terms stay anchored to stable identifiers (provenance + alignment). * **Usage-based sharding:** vocabulary is split into **domain shards** so the engine can load only what it needs for the current context (instead of loading “everything”). * **Strict validation:** every entry must follow a schema so generation doesn’t fail because a required property (e.g., grammatical gender) is missing. ## How it’s organized (human-friendly mental model) * **One namespace per language**, using **ISO 639-1 two-letter codes**. * Inside each language, the Lexicon is split into a few **semantic domains** (files) that match real generation needs: * **core**: “skeleton” function words needed to build any sentence (highest priority). * **people**: terms used for biographies (professions, relations, titles). * **geography**: countries/places and derived forms (adjectives, demonyms). * **science**: specialized terminology (grows over time). ## Why this matters operationally (readiness scoring) SemantiK Architect treats Lexicon coverage as a measurable readiness signal (“Zone B”): a scanner counts words in these shards to grade whether a language is **data-ready** (from “no files” to “production-ready”). ## Typical workflows (high level) * **Bootstrap a new language:** start with **core** first (so basic sentences are possible), then add **people/geography** for biographies. * **Grow coverage from Wikidata:** use Wikidata as the upstream source for translations and QID provenance, then store locally in the right domain shard. * **Fix “missing word” issues:** add the missing term to the appropriate shard (often **people** for biographies), keeping schema-valid entries. ================================================================================================ FILE: wiki/Outputs-Text.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 2f71e2a6c8dbc9671ff1b4f6ab1d8084356f77655976ce04e6332d644b6ba231 CONTENT_BYTES: 1788 ================================================================================================ # Outputs: Text ## What “Text” output is SemantiK Architect’s primary output is **natural language surface text**: the human-readable sentence(s) you would display in a UI, store, or publish. In the engine architecture, the Renderer’s output port is explicitly **“Natural Language Text”** (alongside an optional UD/CoNLL-U output). This output is intended to be **encyclopedic**, and the system is framed around producing **verifiable, high-quality encyclopedic text** across **300+ languages**. ## What you receive (conceptual response envelope) When you request generation, the response includes: * **`surface_text`**: the generated text string (the thing you display) * **`meta`**: minimal provenance about how that string was produced (engine/adapter/strategy), useful for debugging and QA—not required for basic usage. ## What shapes the text output (high level) ### 1) Grammar realization (quality when available) Where strong grammars exist, the system relies on rule-based realization (GF linearization into concrete strings) to produce well-formed text. ### 2) Coverage strategy (consistent output across many languages) SemantiK Architect is designed to keep producing usable text even when a language has limited handcrafted grammar support (via its tiered/hybrid approach). ### 3) Context (when generating more than one sentence) If you generate text in a session, the Context layer can influence surface text choices (e.g., avoiding repetition via pronouns), because it exists specifically to manage discourse state. ## Relationship to UD output (optional) Text is the “publishable” output; UD/CoNLL-U is an **optional companion output** meant to support verification and evaluation. The system treats both as first-class output ports. ================================================================================================ FILE: wiki/Outputs-UD.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: e95fa0ad10717147ef8decfd0edec3f507c44525794e17804c7cce1177489abf CONTENT_BYTES: 1636 ================================================================================================ # 14. Outputs: UD SemantiK Architect can optionally output **Universal Dependencies (UD)** as a **CoNLL-U-style** representation of the generated sentence, alongside the normal surface text. ## What this output is * A **dependency-grammar view** of the same sentence you just generated (token-level structure and relations), produced by the **UD Exporter / UD Mapping** component. * Conceptually, it’s an **“audit trail”**: a standardized linguistic structure you can compare across languages and against UD resources. ## Why SemantiK Architect outputs UD * **Evaluation & benchmarking:** the docs explicitly frame UD output as something you can evaluate “against treebanks.” * **Interoperability:** UD is treated as a standards layer (“Standards: UD Exporter (CoNLL-U Tag Mapping)”). * **Separation of concerns:** UD Mapping is a first-class *output port* alongside the text renderer (so you can use it without changing the core engine). ## When you should use UD output * When you want to **validate** that generation follows a consistent syntactic structure across languages. * When you want a **language-agnostic diagnostic** view (useful when surface text quality varies due to tier/coverage). * When building QA workflows where “text looks OK” is not enough and you want a structural signal. ## Important constraints (high level) * UD output only works as well as the **mapping coverage**: every syntactic constructor needs a corresponding CoNLL-U mapping rule. * If a mapping is missing, UD export can fail (the API error table explicitly calls out “UD Exporter failed to map a function”). ================================================================================================ FILE: wiki/Positioning-Ninai-Udiron-GF-UD-No-WMF-AW-Affiliation.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 44ae0cf3d25948a06346f991b26ab1e66f2fb32e45c7f117757f4d2cc32833a0 CONTENT_BYTES: 3332 ================================================================================================ # Positioning: Ninai, Udiron, GF, UD (No WMF/AW Affiliation) SemantiK Architect (formerly “Abstract Wiki Architect”) is an **independent project**. It is **not affiliated with WMF / Abstract Wiki**, and does not imply collaboration or endorsement. ## The short version (one paragraph) SemantiK Architect is a **renderer platform**: it takes a **language-independent meaning representation** and produces **encyclopedic text** across many languages, optimizing for the “long tail” (high coverage) while staying **deterministic** and **verifiable**. ## Who does what (clear boundaries) ### Ninai — meaning (input standard) * **Role:** A language-independent way to express meaning as constructor trees (“what you want to say”). * **What it’s not:** It does not decide word order, morphology, or surface realization. * **Relationship to SemantiK Architect:** SemantiK Architect can **consume Ninai-style object trees** as an input path. ### GF — high-quality grammar (best-case realization) * **Role:** A rule-based grammar system (abstract syntax → concrete syntax) that can produce high-quality text where strong grammars exist. * **What it’s not:** It is not a full product platform (language selection policy, lexicon lifecycle, QA strategy, discourse memory). * **Relationship to SemantiK Architect:** SemantiK Architect uses GF as a **Tier-1 “high road”** option where available, inside a broader system that must still cover under-resourced languages. ### Udiron — topology/ordering approach (coverage accelerator) * **Role (in this doc set):** A dependency/topology-inspired approach to **linearize** constituents via configurable ordering (weights), used to scale faster than hand-writing full grammars. * **What it’s not:** It is not a complete end-to-end renderer stack (lexicon + tiering + API + context + QA). * **Relationship to SemantiK Architect:** SemantiK Architect adapts a **weighted topology** strategy as the coverage fallback (Tier-3). ### UD (Universal Dependencies) — evaluation/interchange surface (validation) * **Role:** A standard way to represent dependency structure, useful for **evaluation** and interoperability. * **What it’s not:** UD is not a generator; it does not produce text by itself. * **Relationship to SemantiK Architect:** SemantiK Architect can **export UD/CoNLL-U** as an output format to make correctness/analysis measurable and comparable. ## What SemantiK Architect adds (the “product layer” the others don’t) SemantiK Architect’s unique responsibility is the **end-to-end orchestration** from meaning → text at scale: * **Hybrid realization strategy:** chooses between high-quality grammars and coverage fallbacks (tiered approach). * **Lexicon + grounding:** manages vocabulary as a first-class system (including grounding to identifiers such as Wikidata QIDs in the current docs). * **Discourse/context memory:** supports multi-sentence coherence (e.g., pronouns) as a core capability. * **Quality workflow:** explicit “gold standard / judge” style validation mindset (beyond “it compiles”). ## Mental model (one line) **Ninai says what you mean → SemantiK Architect decides how to realize it → GF (when strong) or topology (when needed) produces the sentence → UD export helps verify what was produced.** ================================================================================================ FILE: wiki/Renderer.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: ac680e013da4d0ed88ab013fc9dbb34e13db1f181db160dcc73c7fd00868b06a CONTENT_BYTES: 1591 ================================================================================================ ## 7. Renderer The **Renderer** is the “assembly” layer of SemantiK Architect: it takes an **abstract intent** (meaning) and produces **readable text**. ### What the Renderer is responsible for - **Turn intent into language output**: the Renderer is where “meaning becomes wording.” - **Support two input styles (dual-path)**: - a **strict/production path** (stable, predictable structured input), and - a **prototype path** that can accept a **recursive Ninai-style meaning tree**. - **Expose two output forms**: - **natural language text**, and optionally - a **Universal Dependencies (CoNLL-U)** representation for validation/evaluation workflows. ### What the Renderer decides (high level) - **Which realization strategy to use** for a given language and input (e.g., “high-resource vs long-tail” behavior). This is the practical place where the system chooses between grammar-driven rendering and coverage-oriented fallback approaches. ### What the Renderer does *not* do - It is **not** the lexicon (it doesn’t define vocabulary). - It is **not** the grammar matrix (it doesn’t author the linguistic rules). - It is **not** the context store (it doesn’t own discourse memory; it uses it). ### Why this layer exists Separating the Renderer makes SemantiK Architect easier to scale: - You can improve **inputs** (add richer meaning formats) without rewriting lexicon/grammar. - You can add **outputs** (like UD export) without changing how text is produced. - You can evolve language strategies while keeping a stable “meaning → text” contract. ================================================================================================ FILE: wiki/Repo-Map.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 912e3383db600ee0ea36fd61c02e0681d2b5f8ebe725bf650cb66c0445861379 CONTENT_BYTES: 3148 ================================================================================================ # Repo Map This map lists the **main repo areas you’ll touch most**, grouped by purpose (not exhaustive). It reflects the **current repo layout**, with notes where naming is **legacy** or where behavior is a **target contract** rather than guaranteed in every branch. ## Top-level (conceptual) * **Backend API (server-side)** The backend is where generation, discovery endpoints, and admin/tooling endpoints live. Historically there have been multiple entrypoints; the intent is to converge on a **single canonical API app** and deprecate duplicates over time. * **Frontend UI (client-side)** The UI lives under `architect_frontend/`. It is expected to behave as a single client against the versioned API (see `architect_frontend/src/lib/api.ts` for the central API wrapper). * **Data (lexicon + configuration)** - Lexicon data is stored per language under `data/lexicon/{LANG}/...`. - `{LANG}` is typically the **two-letter ISO code** for internal directories (e.g., `data/lexicon/en/`). If you also maintain external/public identifiers, treat them as a **mapped surface**, not a directory rule. - Key mappings and configuration live under `data/config/` (e.g., `data/config/iso_to_wiki.json`). * **Grammars (grammar sources + compiled artifact)** Grammar sources live under `gf/` (abstract syntax + per-language grammars such as `gf/Wiki{WikiCode}.gf`). A compiled **PGF** artifact is typically generated under `gf/` as part of the build. Note: some filenames may still carry **legacy naming** from the project’s earlier name. * **Generated sources (build outputs)** Generated language sources are produced under `generated/`. If there are multiple mirrors (e.g., an older `gf/generated/...` layout), treat `generated/` as the **preferred** source of truth and regard mirrors as legacy compatibility. * **Tools & utilities (operators + QA + data ops)** - `tools/` contains operator-facing tools (including QA suites). - `utils/` contains helper scripts and data operations. This split may evolve; avoid hardcoding paths in automation where possible. The Tools Registry is served by the backend router (commonly referenced under `app/adapters/api/routers/...`). * **Deployment** Deployment wiring typically lives in `deploy/` (reverse proxy config, etc.) and `docker-compose.yml` (services wiring). * **Project launcher / orchestrator** `manage.py` is the main “project command” entrypoint used for common flows (build/clean/align), depending on how your repo is operated. --- ## Suggested tree view (high level) * `app/` — Backend API (canonical entrypoint under `app/adapters/api/...`) * `architect_frontend/` — Frontend UI * `data/` * `data/lexicon/{LANG}/...` — Lexicon per language * `data/config/...` — Configuration + mappings (e.g., `iso_to_wiki.json`) * `gf/` — Grammar sources + compiled grammar artifact (PGF) * `generated/` — Preferred generated sources * `tools/` — Operator tools / QA * `utils/` — Data ops + helper scripts * `deploy/` — Reverse proxy configuration * `docker-compose.yml` — Services wiring * `manage.py` — Orchestration entrypoint ================================================================================================ FILE: wiki/Roadmap.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: bcffae98554b68e72ba6089fd2a4f94b35546bfe7f7e75e7282bdc2b106b9bad CONTENT_BYTES: 2588 ================================================================================================ # Roadmap — SemantiK Architect > Formerly “Abstract Wiki Architect”. SemantiK Architect is **independent** and **not affiliated with WMF / Abstract Wiki**. ## Guiding principles - **Deterministic by default**: normal builds and runtime operation do not rely on automated LLM calls. - **Human-in-the-loop for hard repairs**: any AI assistance (if enabled) is interactive and reviewable, not an unattended build hook. - **Quality is enforceable**: improvements must be measurable and protected against regressions (Gold Standard + gating). --- ## Near-term (Wiki v1 scope) ### 1) Product clarity + documentation stabilization - Publish the **short GitHub Wiki** (Home, Positioning, How it Works, Quality, Using, Dev Notes, Project). - Align wording across pages on AI usage and determinism (remove contradictions, mark “target contract” vs “current behavior”). - Normalize naming (“SemantiK Architect”) across UI/docs/repo, and clearly label legacy repo paths/names where they still exist. ### 2) Language onboarding (reliable first pass) - Make “Add a language” consistently result in a **runnable** language (minimum lexicon seed + basic grammar selection). - Keep the **Everything Matrix** as the single discovery/health registry that drives the language strategy (tier selection, readiness signals). ### 3) Verifiability baseline - Establish a **Gold Standard** set for high-signal sentence types and keep gating behavior stable and documented. - Keep **UD export** usable as a validation surface where applicable. --- ## Execution order (implementation roadmap, simplified) 1. **Foundation**: naming cleanup, configuration cleanup, and a clear “language strategy” registry model. 2. **Deterministic adapters**: stable input adapters (Frames + Ninai) and stable validation/export surfaces (e.g., UD). 3. **Core generation coverage**: ensure Tier 3/fallback coverage is reliable; integrate tier selection cleanly. 4. **Context support**: enable multi-sentence coherence in a controlled, testable way. 5. **Optional AI tools**: keep AI-assisted workflows out of the default build path; improve evaluation tooling and developer experience. --- ## Later (v2 — explicitly out of “simple wiki” scope) - **Learned micro-planning**: optional pre-render rewriting for style variation (tone, synonym choice) before deterministic rendering. - **Richer discourse planning**: broader context control beyond basic reference/pronouns. - **Expanded validation**: broader gold standards, stronger automated regression analysis, and clearer quality dashboards. ================================================================================================ FILE: wiki/Setup.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 524dbcfdf2691d5601a4fc1727d7ed3525b532b7d584735428a2840b6ac49323 CONTENT_BYTES: 2736 ================================================================================================ ## 17. Setup This page explains how to get **SemantiK Architect** running locally (for development) and how to deploy it (for production). The system’s backend runtime depends on **GF C-libraries**, so it must run in a **Linux environment**; on Windows, the recommended approach is a **Windows + WSL2 hybrid** (Windows for editing/frontend, WSL2 for backend + GF). ### Supported development setup (recommended) * **Windows 10/11 + WSL2 (Ubuntu)**: edit code and run the frontend on Windows, run backend services and GF tooling in WSL2. * Keep the repo on the **Windows drive** so both Windows and WSL can access the same working tree; avoid cloning into the Linux-only filesystem if you edit in Windows. ### Prerequisites (high level) You’ll need: * WSL2 + Ubuntu (22.04+ suggested in the docs) * Docker Desktop (WSL2 backend enabled) * VS Code with the WSL extension * Node.js 18+ (for the frontend) ### Core setup steps (conceptual) 1. **System initialization (Linux side):** install OS build dependencies and the GF toolchain, and obtain the GF Resource Grammar Library (RGL). 2. **Python environment (Linux side):** create a virtual environment and install Python dependencies. 3. **Configuration:** create a `.env` at repo root to point the backend to repo paths, GF/RGL paths, compiled grammar artifact path, Redis, and tool limits. AI-related tools are explicitly gated by an environment flag. ### Running locally The docs describe two ways: * **Unified launcher (recommended):** one entry point that starts API, worker, and frontend. * **Manual mode:** run Redis, the API backend, the worker, and the frontend as separate processes. The UI is served under a **base path** (not at the root), and there are dedicated dev/tools pages. ### Production deployment (Docker) For production/full-stack container runs, the docs use `docker-compose`. Key conceptual differences vs local: * repo path is mounted into containers * services talk via container hostnames (not `localhost`) * UI and API are served under the configured base path and versioned API prefix ### Troubleshooting (common themes) * Backend errors about **pgf/lib** typically mean you’re running outside the Linux environment or missing system deps. * Windows/WSL file **line endings** can break scripts; convert when needed. * API calls should use the **versioned `/api/v1` prefix**. * Tool execution failures usually come from an incorrect repo-root path configuration. * AI tools may return authorization/gating errors if they’re disabled in configuration. ### Verification (smoke test) The docs include two end-to-end checks: * a **standard structured frame** generation request * a **Ninai protocol** generation request ================================================================================================ FILE: wiki/What-SemantiK-Architect-Is.md AUTHORITY: reference CONTENT_ROLE: navigation CONTENT_SHA256: 87322206668f233d39f5dab3aadec45f4aee720b81a1ddfc56b3d299a155216b CONTENT_BYTES: 2078 ================================================================================================ ## 2. What SemantiK Architect Is **SemantiK Architect** (renamed from “Abstract Wiki Architect” in these docs) is an **independent** project: a **hybrid Natural Language Generation (NLG) engine** designed to produce **verifiable encyclopedic text** in the **long tail of languages (300+)**. ### What it does (in one sentence) It takes a **language-independent intent** and turns it into **natural language text**, with an emphasis on **determinism**, **scalability across many languages**, and **evaluation-ready outputs**. ### The guiding idea SemantiK Architect combines: * A **deterministic, rule-based core** for correctness, and * **AI-assisted tools at the edges** (copilot/QA/bootstrapping) to scale language coverage without turning generation into free-form “chat text”. ### The four core “building blocks” (conceptual) The system is organized as four layers, so each concern stays clear and replaceable: 1. **Lexicon**: words + linguistic properties, grounded to **Wikidata QIDs** for semantic consistency. 2. **Grammar**: rules for word forms and word order, with a hybrid approach for both high-resource and under-resourced languages. 3. **Renderer**: the assembly step that converts intent into text and supports multiple input styles. 4. **Context**: a memory layer to support multi-sentence behavior (e.g., pronouns, discourse continuity). ### Inputs and outputs (high level) * **Inputs:** * A **strict** structured “frame” for stable production behavior, and * A **recursive Ninai-style object tree** for expressive/experimental meaning input. * **Outputs:** * **Surface text**, and optionally * A **Universal Dependencies (CoNLL-U)** representation intended for evaluation/validation workflows. ### What it is *not* * Not a knowledge base or encyclopedia: it’s the **rendering engine**, not the source of truth. * Not a “prompt-and-pray” text generator: it is explicitly **generation-first and structured**. * Not affiliated with the WMF/Abstract Wiki team (project note based on your stated context). ================================================================================================ FILE: data/raw_wikidata/README.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: e0b8e962a227a6db71b7be9fed1961b6c660d7672118c5d65d4ddb4bbb904bc5 CONTENT_BYTES: 1511 ================================================================================================ # data/raw_wikidata/ This directory is reserved for **local Wikidata / Lexeme dump files** used to build or refresh the project lexica. These files are **not** meant to be committed to version control (see `.gitignore` in this directory). ## What goes here? Examples of files you might place here: - `lexemes_dump.json.gz` A filtered Wikidata Lexeme dump (e.g. only selected languages / domains). - `entities_dump.json.gz` A filtered dump of Q-items (e.g. people, countries) used to enrich lexicon entries with `qid`, `country_qid`, etc. File names are not fixed; the build scripts in `utils/` should take the actual paths as CLI arguments. ## How is this used? Typical workflow: 1. Download or prepare a Wikidata / Lexeme dump. 2. Put the dump in this directory (or somewhere on disk). 3. Run the builder script, for example: ```bash python utils/build_lexicon_from_wikidata.py \ --lang it \ --dump data/raw_wikidata/lexemes_dump.json.gz \ --out data/lexicon/it_lexicon.json ```` 4. Run QA tools on the resulting lexicon: ```bash python qa_tools/lexicon_coverage_report.py ``` ## Git policy * Raw dumps can be **very large** and change often. * They should **never** be committed to the repository. * Only keep: * this `README.md`, * the `.gitignore` file, * very small synthetic fixtures if needed for tests. If you need to share a specific dump configuration, document the download command or filtering pipeline instead of the file itself. ================================================================================================ FILE: data/reports/lexicon_coverage_report.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 4d83847cc10e65a82f99b66b5d1f339d5dedc859f6ce44fc8c0d6b06c3c88aa9 CONTENT_BYTES: 591 ================================================================================================ # Lexicon Coverage Report - Generated at: `2026-03-26T18:03:11Z` - Lexicon dir: `/mnt/c/mycode/SemantiK_Architect/SemantiK_Architect/data/lexicon` - Targets: core=150, conc=500, bio_min=50 ## Summary | lang | core | conc | qids | lexemes | collisions | errors | warnings | CORE | CONC | BIO | | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | | fr | 55 | 18 | 4 | 73 | 2 | 0 | 0 | 3.67 | 0.36 | 0.80 | ## Totals ```json { "languages": 1, "sum_lexemes": 73, "sum_qids": 4, "sum_collisions": 2, "sum_errors": 0, "sum_warnings": 0, "missing_core": [] } ``` ================================================================================================ FILE: docs/Technical-Reference/00-SETUP_AND_DEPLOYMENT.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: d6251a1b2846a1784636c11c9f36a32bbd80add20c39fb6ae6f8326fa85b7de5 CONTENT_BYTES: 10257 ================================================================================================ # 🛠️ Setup & Deployment Guide **SemantiK Architect — current repository layout** This guide covers installation, configuration, local development, and Docker deployment for the current repository. Because the backend depends on **Grammatical Framework (GF)** and the Python `pgf` bindings, the backend must run in a **Linux environment**. On Windows, the recommended setup is hybrid: 1. **Windows 11**: editing, Git, frontend 2. **WSL 2 (Ubuntu)**: backend, Python environment, Redis, GF --- ## 1. Prerequisites Recommended host setup: - **Windows 10/11** with **WSL 2** - **Ubuntu** installed in WSL - **VS Code** with the **WSL** extension - **Node.js 18+** on Windows for the frontend - **Docker Desktop** if you want Redis via Docker and/or full-stack Docker deployment Inside WSL, install the base Linux packages you need for Python builds and GF-related tooling: ```bash sudo apt update sudo apt install -y \ python3-venv python3-dev \ build-essential libgmp-dev \ git curl wget dos2unix ```` > Keep the repository on your Windows drive if you want both Windows and WSL to share the same working tree. --- ## 2. Repository Layout Use the repository root that contains `manage.py`, `pyproject.toml`, `requirements.txt`, `app/`, and `architect_frontend/`. Recommended layout: ```text C:\MyCode\SemantiK_Architect\ └── SemantiK_Architect\ <-- repo root ├── .env ├── .venv/ <-- local Python environment ├── manage.py ├── pyproject.toml ├── requirements.txt ├── app/ ├── architect_frontend/ ├── builder/ ├── data/ ├── docs/ ├── gf/ │ └── semantik_architect.pgf ├── gf-rgl/ <-- clone or place gf-rgl here ├── schemas/ └── tools/ ``` In WSL, the same repo will usually be visible as: ```bash /mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect ``` > `gf-rgl/` is expected inside the repo root for the current toolchain and Docker setup. --- ## 3. GF / RGL Setup (WSL) Open a **WSL terminal** and go to the repo root: ```bash cd /mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect ``` ### A. Install GF Install a GF package appropriate for your Ubuntu release, then verify: ```bash gf --version ``` If `gf` is already available in your WSL environment, you can skip this step. ### B. Ensure `gf-rgl/` exists in the repo root If `gf-rgl/` is not already present: ```bash git clone https://github.com/GrammaticalFramework/gf-rgl.git ``` This should create: ```text /gf-rgl/ ``` --- ## 4. Python Environment (WSL) For the current repository state, keep local installation simple: **one Python environment at the repo root**. Preferred: ```bash uv venv .venv source .venv/bin/activate uv pip install -r requirements.txt ``` Fallback if you are not using `uv`: ```bash python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` > For now, install from `requirements.txt`. The repo also contains a `pyproject.toml`, but the current worker Docker image still installs from `requirements.txt`, so this is the least surprising local setup. --- ## 5. Environment Configuration (`.env`) Create a `.env` file in the repo root. **File:** `.env` ```ini # --- App --- APP_NAME=Semantik Architect APP_ENV=development DEBUG=true # --- Repo / Filesystem --- FILESYSTEM_REPO_PATH=/mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect # --- GF / PGF --- GF_LIB_PATH=/mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect/gf-rgl PGF_PATH=/mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect/gf/semantik_architect.pgf # Legacy compatibility only (optional) AW_PGF_PATH=/mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect/gf/semantik_architect.pgf # --- Redis / Worker --- REDIS_URL=redis://localhost:6379/0 REDIS_QUEUE_NAME=architect_tasks # --- Local browser dev --- ARCHITECT_CORS_ORIGINS=http://localhost:3000,http://127.0.0.1:3000 # --- Optional: only set this when backend is mounted behind a reverse proxy --- # ARCHITECT_API_ROOT_PATH=/semantik_architect # --- Optional API auth --- # API_KEY=change-me # API_SECRET=change-me ``` Notes: * Prefer **`PGF_PATH`**. `AW_PGF_PATH` is legacy compatibility. * `FILESYSTEM_REPO_PATH` should point to the repo root visible from the backend runtime. * `ARCHITECT_API_ROOT_PATH` is only needed when the backend is mounted behind a prefix such as `/semantik_architect`. --- ## 6. Local Development There are three practical ways to run the stack. ### Option A — Windows launcher (recommended on Windows) From **PowerShell** at the repo root: ```powershell .\Run-Architect.ps1 ``` What it does: * starts the API in WSL * starts the worker in WSL * starts the frontend on Windows * probes backend readiness * opens the default tools UI URL Useful flags: ```powershell .\Run-Architect.ps1 -NoPortClear .\Run-Architect.ps1 -EnableReload -EnableWatch .\Run-Architect.ps1 -SkipFrontend ``` ### Option B — Canonical CLI orchestrator (WSL) From WSL at the repo root: ```bash source .venv/bin/activate python manage.py doctor python manage.py start ``` Useful commands: ```bash python manage.py build --align python manage.py clean python manage.py doctor ``` ### Option C — Manual run (4 terminals) #### Terminal 1 — Redis With Docker: ```bash docker run -p 6379:6379 --name aw_redis -d redis:alpine ``` #### Terminal 2 — API backend (WSL) ```bash cd /mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect source .venv/bin/activate python -m uvicorn app.adapters.api.main:create_app --factory --host 0.0.0.0 --port 8000 --reload ``` #### Terminal 3 — Async worker (WSL) ```bash cd /mnt/c/MyCode/SemantiK_Architect/SemantiK_Architect source .venv/bin/activate python -m arq app.workers.worker.WorkerSettings --watch app ``` #### Terminal 4 — Frontend (Windows PowerShell) ```powershell cd architect_frontend npm install npm run dev ``` ### Local URLs Frontend: * `http://localhost:3000/semantik_architect` * `http://localhost:3000/semantik_architect/dev` * `http://localhost:3000/semantik_architect/tools` Backend: * `http://localhost:8000/docs` * `http://localhost:8000/health/ready` * `http://localhost:8000/api/v1/health/ready` --- ## 7. API Notes Current backend behavior: * Canonical API prefix: **`/api/v1`** * Local dev frontend expects backend on `:8000` * Local dev frontend origin `:3000` is allowed by default in development * The generate endpoint is available at **`POST /api/v1/generate/{lang_code}`** * A payload-based variant also exists at **`POST /api/v1/generate`** For the safest smoke tests, use the language-path form: ```bash curl -X POST "http://localhost:8000/api/v1/generate/en" \ -H "Content-Type: application/json" \ -H "X-API-Key: ${API_KEY}" \ -d '{ "frame_type": "bio", "subject": { "name": "Marie Curie", "qid": "Q7186" }, "properties": { "label": "Marie Curie" } }' ``` If your local environment does **not** enforce API auth, you can omit `X-API-Key`. Expected response shape: ```json { "text": "…", "lang_code": "en" } ``` --- ## 8. Docker Deployment The repo includes: * `docker/Dockerfile.backend` * `docker/Dockerfile.worker` * `docker/Dockerfile.frontend` * `docker-compose.yml` * `deploy/nginx.conf` ### Full stack Run from the repo root: ```bash docker-compose up --build -d ``` ### What starts * `redis` * `backend` * `worker` * `frontend` * `nginx` ### Public URLs The reverse proxy publishes the app here: * `http://localhost:4000/semantik_architect/` * `http://localhost:4000/semantik_architect/tools` * `http://localhost:4000/semantik_architect/api/v1/...` ### Docker notes * Inside containers, the repo is mounted at `/app` * The frontend is served under `/semantik_architect` * Nginx proxies `/semantik_architect/api/*` to the backend * Docker currently uses: * backend on port `8000` * frontend on port `3000` * nginx on host port `4000` --- ## 9. Troubleshooting ### `pgf` module not found Cause: * backend running on Windows instead of Linux/WSL * missing build prerequisites such as `libgmp-dev` Fix: ```bash sudo apt install -y libgmp-dev build-essential python3-dev source .venv/bin/activate pip install -r requirements.txt ``` ### `gf-rgl/ folder missing` Cause: * repo root does not contain `gf-rgl/` Fix: ```bash cd git clone https://github.com/GrammaticalFramework/gf-rgl.git ``` ### Frontend says `Failed to fetch` Cause: * backend is not running * CORS is misconfigured * frontend is trying to hit the wrong API base Fix: * confirm backend on `http://localhost:8000` * confirm `ARCHITECT_CORS_ORIGINS` includes `http://localhost:3000` * confirm you are using the canonical `/api/v1` routes ### `404 Not Found` on the API Cause: * missing `/api/v1` prefix * wrong base path under Docker/reverse proxy Fix: * local dev backend: use `/api/v1/...` * behind Nginx: use `/semantik_architect/api/v1/...` ### Build artifacts end up in the wrong place Cause: * generated GF files or PGF are being written outside the expected repo locations Fix: * keep generated grammar artifacts under `gf/` * keep `PGF_PATH` pointing to `gf/semantik_architect.pgf` --- ## 10. Verification Checklist From WSL: ```bash source .venv/bin/activate python manage.py doctor ``` Then verify: 1. `gf --version` works 2. `gf-rgl/` exists at repo root 3. `gf/semantik_architect.pgf` exists after a build 4. `http://localhost:8000/api/v1/health/ready` returns 200 5. `http://localhost:3000/semantik_architect/tools` loads 6. `curl` to `/api/v1/generate/en` returns 200 or a clear auth/validation response --- ## 11. Minimal Day-to-Day Commands ### Start everything (Windows hybrid) ```powershell .\Run-Architect.ps1 ``` ### Start everything from WSL ```bash source .venv/bin/activate python manage.py start ``` ### Rebuild grammar layer ```bash source .venv/bin/activate python manage.py build --align ``` ### Frontend only ```powershell cd architect_frontend npm run dev ``` ### Backend only ```bash source .venv/bin/activate python -m uvicorn app.adapters.api.main:create_app --factory --host 0.0.0.0 --port 8000 --reload ``` ================================================================================================ FILE: docs/Technical-Reference/01-ENGINE_ARCHITECTURE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 3686215728cdefa5d034a5ef462136fa67996e5125a91f09e861135d75e7e73c CONTENT_BYTES: 15325 ================================================================================================ # 🏛️ Engine Architecture & Internals **SemantiK Architect v2.1** ## 1. High-Level System Overview SemantiK Architect is a **planner-centered multilingual NLG system** for generating structured, traceable text from semantic inputs. Its architectural center is **not** any one renderer, grammar formalism, or language-specific engine. The stable center is the shared runtime pipeline: ```text API/request -> frame normalization -> frame-to-plan bridge -> planner -> PlannedSentence -> ConstructionPlan -> lexical resolution -> renderer backend -> SurfaceResult -> API response mapping ```` This design allows SemantiK Architect to combine: * **semantic frames** as structured input, * **construction planning** as the source of sentence truth, * **lexical resolution** as a shared multilingual layer, * **multiple realization backends** such as GF/PGF, family renderers, and safe-mode fallback. The system is built for **deterministic, inspectable multilingual generation**, with explicit support for broad language coverage and backend diversity. --- ## 2. Architectural Principles ### 2.1 One source of runtime truth The planner and shared construction runtime define **what is being said**. Renderers define only **how that plan is realized** in a particular backend or language. No renderer, engine, or router should be an independent source of sentence-planning truth. ### 2.2 Construction-first, not bio-first Biography is an important early domain and migration target, but it is **not** the architecture. The runtime is intended to support multiple construction families, including: * equative / classification * attributive copular * locative * existential * possession * eventive * relative-clause * topic-comment * comparative / superlative * coordination ### 2.3 Backend independence The same semantic intent should be able to flow through different realization technologies: * **GF / PGF** * **family renderers** * **safe-mode fallback** GF is a renderer backend and tooling source, **not the architecture itself**. ### 2.4 Shared semantics, thin language specialization The architecture scales by sharing: * frame normalization, * planning, * construction IDs, * slot semantics, * lexical binding structure, * renderer contracts. Language-specific logic should be limited to: * lexical forms, * morphology, * local syntax / word order, * idiomatic overrides, * construction-specific realization details where required. --- ## 3. Runtime Layers ### Layer A: Semantic Inputs & Frame Normalization **Role:** Convert external inputs into stable internal generation commands. This layer accepts API payloads and normalizes them into domain objects. Current input families include: * flat frame payloads such as biography/person-style requests, * compatibility aliases for person/bio payloads, * Ninai-shaped payloads where supported. Key responsibilities: * resolve the authoritative language code, * normalize payload variants, * reject malformed or contradictory requests, * strip transport-only fields from the domain payload. This layer is where input compatibility is handled. It is **not** where sentence planning should live. --- ### Layer B: Planning & Construction Runtime **Role:** Decide the sentence structure to be realized. This is the architectural core. The planner-centered runtime is responsible for producing a shared representation of sentence intent, including: * `PlannedSentence` * `construction_id` * topic/focus metadata * slot layout * realization options * construction-level semantics This layer is the authoritative center for: * sentence type, * information packaging, * construction choice, * discourse-aware decisions. The target runtime contract is: ```text frame normalization -> frame-to-plan bridge -> planner -> PlannedSentence -> ConstructionPlan ``` --- ### Layer C: Lexicon & Lexical Resolution **Role:** Bind planned slots to language-appropriate lexical material. Lexical resolution is a distinct layer between planning and realization. It is not just preprocessing. This layer handles: * entity references, * lexical bindings, * lemma selection, * morphology-relevant lexical metadata, * provenance such as Wikidata QIDs or source-specific IDs where available. The lexicon remains a separate subsystem and a shared concern across backends. This separation is important because the same construction plan may be realized by different backends, but they must operate over the same lexicalized intent. --- ### Layer D: Renderer Backends & Surface Realization **Role:** Realize a shared construction-level plan into surface output. Renderer backends are interchangeable surface technologies operating over the same runtime contract. Supported backend classes include: * **GF / PGF renderer** * **family-oriented renderer** * **safe-mode fallback renderer** The renderer produces a `SurfaceResult`, which is then mapped to the public API response. Possible output forms include: * natural-language text, * debug traces / runtime metadata, * backend-specific diagnostics, * structured exports where supported. A backend may be strong for some `(lang_code, construction_id)` pairs and unavailable for others. Capability is therefore backend-specific, language-specific, and construction-specific. --- ### Cross-Cutting Concern: Context & Discourse State **Role:** Support discourse-aware generation beyond isolated single sentences. The repository includes discourse planning and stateful components for things such as: * topic tracking, * focus management, * referring expressions, * pronominalization, * session continuity. This is a **cross-cutting runtime concern**, not a separate renderer. When enabled, context should influence planning and reference choice through the planner/runtime layer, not by ad hoc string rewriting after realization. --- ## 4. Current Runtime Status vs Target Runtime ### Target architecture The intended authoritative runtime is: ```text API payload -> frame normalization -> frame-to-plan bridge -> planner -> PlannedSentence -> ConstructionPlan -> lexical resolution -> renderer backend -> SurfaceResult -> API response mapping ``` ### Current live behavior The repository still documents and supports a **compatibility path** in which single-sentence generation may bypass the full planner-centered path and call a realization engine more directly after frame normalization. That compatibility path is operational and useful during migration, but it is **not** the final architectural center. So the correct reading is: * **planner-centered construction runtime** = target source of truth * **direct frame-to-engine generation** = compatibility shim during migration This distinction matters because “it generates text” is not the same as “it is fully aligned with the final runtime contract.” --- ## 5. Realization Strategy & Language Coverage SemantiK Architect supports a hybrid realization strategy to balance quality, scale, and graceful degradation. ### Tier 1: GF / expert-grade realization Used where strong concrete grammars and runtime support exist. Strengths: * richer morphology, * stronger syntax control, * higher-quality realization for supported constructions. ### Tier 2: Curated / override layers Used where manual or project-maintained overrides improve quality beyond generic defaults. This layer can refine language-specific behavior without redefining the architecture. ### Tier 3: Safe-mode / fallback realization Used to preserve runtime continuity when stronger renderers are unavailable for a given language/construction pair. This layer exists for coverage and fault tolerance, not as the architectural ideal. ### Important rule A language being routable to a backend is **not** the same thing as that language being fully validated for production-grade realization. Validation must be tracked separately at the level of: * language, * construction family, * backend, * test coverage, * gold-example quality. --- ## 6. GF in the Architecture GF is an important part of the system, but its role must be understood correctly. ### GF is: * a realization backend for selected runtime plans, * an offline source of grammar knowledge and QA examples, * a valuable resource for high-quality multilingual realization. ### GF is not: * the core architecture, * the planner, * the public API contract, * the authoritative semantic representation, * the only supported realization technology. SemantiK Architect may compile and use GF assets such as: * abstract grammar definitions, * concrete `Wiki*` modules, * PGF artifacts, * language-specific GF-backed realization paths. But runtime authority remains with the shared planner/construction contract. --- ## 7. Build, Artifacts, and Validation The build/tooling layer exists to assemble, validate, and audit multilingual realization assets. Important concerns include: * language inventory / matrix generation, * grammar-path discovery, * compile audits, * runtime health checks, * capability tracking, * language-level validation. The key artifact for GF-backed runtime use is the compiled PGF/grammar set, but architecture correctness cannot be inferred from build success alone. A language is only meaningfully integrated when the system can demonstrate: * successful build or capability discovery, * successful runtime generation, * correct construction-level behavior, * acceptable surface quality, * regression-safe validation. --- ## 8. Hexagonal Architecture The backend follows **ports and adapters** so that domain logic stays isolated from infrastructure. ### Core domain / application responsibilities The core is responsible for: * frame/domain models, * planning logic, * construction-level abstractions, * shared runtime contracts, * use-case orchestration. ### Adapters Adapters connect the core to external systems, including: * the HTTP API, * persistence and filesystem access, * GF runtime wrappers, * Redis or messaging infrastructure, * exporter layers, * tooling and operational services. Dependencies point inward: adapters depend on core/application code, not the reverse. This structure keeps the architectural center stable even as external technologies change. --- ## 9. Automation, QA, and Operational Tooling The repository contains optional tooling and agent-oriented components for authoring, repair, QA, and maintenance. These may include builder, judge, or repair-oriented workflows. They should be understood as **operational tooling**, not as part of the deterministic core runtime contract. Core generation should remain understandable without assuming any AI service is active. The right separation is: * **core runtime** = deterministic generation architecture * **automation/tooling** = assistance for build, QA, repair, or maintenance --- ## 10. Request Lifecycle ### 1. Ingest A client sends a request to the generation API. Typical route shape: ```text POST /api/v1/generate/{lang_code} ``` ### 2. Normalize The API layer: * resolves the authoritative language code, * validates payload shape, * normalizes compatible frame variants, * maps the request into domain form. ### 3. Plan The target path converts the normalized frame into a plan-oriented representation and produces: * `PlannedSentence` * `ConstructionPlan` * runtime metadata such as `construction_id` ### 4. Resolve Lexicon The system binds planned slots to lexical material needed for realization. ### 5. Realize A selected backend realizes the construction plan: * GF / PGF if available and appropriate, * family renderer if selected, * safe mode if needed. ### 6. Map to Public Response The internal surface result is converted into the public API response. The public response should remain clearly distinguished from internal runtime objects. --- ## 11. Public API Response vs Internal Runtime Objects It is important not to confuse internal and external contracts. ### Internal runtime objects Examples include: * `PlannedSentence` * `ConstructionPlan` * lexical bindings * `SurfaceResult` These are runtime contracts inside the system. ### Public API response The API exposes a user-facing response envelope, including surface text and runtime/debug metadata intended for clients. The public API response is a mapped external view of internal runtime results. It should not be documented as if it were the runtime contract itself. --- ## 12. Directory Map & Key Files | Path | Component | Role | | ---------------------------------------------------------- | ---------------------- | ------------------------------------------------------------------------------- | | `app/adapters/api/contracts/generation_request_mapper.py` | Request normalization | Resolves language, normalizes payload variants, maps HTTP input to domain input | | `app/adapters/api/contracts/generation_response_mapper.py` | Response mapping | Maps internal generation results to the public API contract | | `app/adapters/api/routers/generation.py` | API route | HTTP entry point for generation | | `app/adapters/engines/` | Renderer backends | GF, family, and safe-mode realization adapters | | `app/adapters/persistence/lexicon/` | Lexical infrastructure | Lexicon access, caching, indexing, and entity/lexeme resolution | | `discourse/` | Context and discourse | State, referring expressions, discourse planning | | `gf/` | GF grammars | Abstract and concrete GF modules used by GF-backed realization | | `schemas/contracts/` | Runtime contracts | Shared schemas for construction/runtime structures | | `schemas/frames/` | Frame schemas | Structured semantic input families | | `tools/language_health/` | Validation tooling | Compile audits, runtime checks, reporting | | `tools/everything_matrix/` | Inventory/tooling | Language/resource scanning and matrix generation | | `builder/orchestrator/` | Build orchestration | Build and assembly workflows for grammar/runtime assets | --- ## 13. What This System Is Not SemantiK Architect is **not**: * a pure template engine, * a GF-only architecture, * a biography-only system, * a renderer-first design, * an LLM-only generation stack, * a system where build success alone proves language readiness. It is best understood as a **planner-first, construction-centered, backend-flexible multilingual NLG platform**. ================================================================================================ FILE: docs/Technical-Reference/02-BUILD_SYSTEM.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: a58cea9ca07ff41936bde886cba6b3a6710657ae34e22c0d95c6b07cb3edb956 CONTENT_BYTES: 7508 ================================================================================================ # 🏗️ The Build System & Everything Matrix **SemantiK Architect v2.1** ## 1. Overview: Data-Driven Orchestration In a traditional system, you might hardcode a list of supported languages (e.g., `LANGS = ['en', 'fr']`). In the SemantiK Architect, this is forbidden. Instead, we use a **Data-Driven Architecture**. The system scans its own file system to discover what languages are available, grades their quality, and dynamically decides how to build them. This "Self-Awareness" is stored in a central registry called the **Everything Matrix**. ### Key Principles * **No Hardcoding:** The build script does not know that "French" exists until the Matrix tells it. * **Graceful Degradation:** If a grammar is incomplete (low score), the system automatically downgrades it to "Safe Mode" (Tier 3), utilizing the **Weighted Topology Factory** to ensure valid output. * **Two-Phase Compilation:** We separate *verification* from *linking* to solve the GF "Last Man Standing" overwriting bug. * **AI Intervention:** If a build fails, the **Architect Agent** is triggered automatically to attempt code repair. --- ## 2. The Source of Truth: `everything_matrix.json` The heart of the build system is a JSON file regenerated daily by the scanners. **Location:** `data/indices/everything_matrix.json` ### The Schema The matrix tracks the "Health" of every language across distinct zones. Note that the primary keys are **ISO 639-1 (2-letter)** codes. ```json { "timestamp": 1709337600, "stats": { "total_languages": 75, "production_ready": 12 }, "languages": { "fr": { "meta": { "iso": "fr", "tier": 1, "origin": "rgl", "folder": "french" }, "blocks": { "rgl_grammar": 10, // Zone A: Grammar Logic "rgl_syntax": 10, // Zone A: Sentence Building "lex_seed": 8, // Zone B: Vocabulary Size "lex_wide": 0 // Zone B: Bulk Import status }, "status": { "build_strategy": "HIGH_ROAD", // Decisions: HIGH_ROAD vs SAFE_MODE "maturity_score": 9.2, "data_ready": true } } } } ``` --- ## 3. The Scanning Suite Before a build occurs, a suite of Python scripts audits the codebase to populate the Matrix. These live in `tools/everything_matrix/`. ### A. The Master Indexer (`build_index.py`) * **Role:** The Conductor. It initializes the scan, calls the sub-scanners, aggregates the scores, and writes the final JSON file. * **Command:** `python tools/everything_matrix/build_index.py` ### B. The Grammar Auditor (`rgl_auditor.py`) * **Role:** Audits **Zone A (The Foundation)**. * **Logic:** It physically scans the `gf-rgl/src` directory for the 5 Pillars of RGL: 1. `Cat` (Category definitions) 2. `Noun` (Morphology) 3. `Grammar` (Structural Core) 4. `Paradigms` (Constructors) 5. `Syntax` (API) * **Scoring:** * **10/10:** All 5 modules exist. (Strategy: `HIGH_ROAD`) * **< 7/10:** Missing critical modules. (Strategy: `SAFE_MODE`) ### C. The Lexicon Scanner (`lexicon_scanner.py`) * **Role:** Audits **Zone B (The Vocabulary)**. * **Logic:** It parses the JSON shards in `data/lexicon/{lang_code}/` (e.g., `data/lexicon/fr/`) to count actual words. * **Scoring:** * **0:** No files. (Status: `data_ready = False`) * **5:** Functional core (< 50 words). * **10:** Production ready (> 200 words + Wide Import). --- ## 4. The Maturity Scale (0-10) Every language is assigned a `maturity_score` based on the audit. | Score | Rating | Meaning | Build Action | | --- | --- | --- | --- | | **0 - 2** | 🔴 **Broken** | Critical files missing. | **Skip.** Do not attempt to build. | | **3 - 5** | 🟡 **Draft** | Auto-generated or incomplete. | **Safe Mode.** Build using **Weighted Topology Factory** (Udiron-style ordering). | | **6 - 7** | 🔵 **Beta** | Manual implementation, potentially buggy. | **Safe Mode.** Use RGL but verify strictly. | | **8 - 9** | 🟢 **Stable** | Full RGL support + Lexicon. | **High Road.** Full optimization. | | **10** | 🌟 **Gold** | Production verified, unit tests pass. | **High Road.** | --- ## 5. The Build Orchestrator (`builder/orchestrator.py`) This script reads the Matrix and executes the compilation. It solves the critical "Last Man Standing" bug using a **Two-Phase Pipeline** augmented by AI. ### Phase 1: Verification (The "Try" Loop) The system iterates through every language in the Matrix. It does **not** link them yet. It runs the compiler in "Check Mode" to generate intermediate `.gfo` files. * **Command:** `gf -batch -c -path ... Wiki{Lang}.gf` * **Purpose:** To verify that the code *can* compile without actually creating the final binary. ### Phase 1.5: The Architect Loop (Auto-Repair) [NEW] If Phase 1 fails for a language: 1. **Capture:** The orchestrator captures the GF compiler error log (e.g., `unknown function 'mkN0'`). 2. **Trigger:** It calls the **Architect Agent** with the broken code and the error. 3. **Patch:** The Agent rewrites the `.gf` file to fix the error (e.g., swapping `mkN0` for `mkN`). 4. **Retry:** The system compiles again. (Max Retries: 3). 5. **Drop:** If it still fails after 3 tries, the language is removed from the build list. ### Phase 2: Linking (The "Make" Shot) Once the list of valid languages is finalized, the orchestrator runs **one single command** to link them all together. * **Command:** `gf -batch -make -path ... semantik_architect.gf WikiEng.gf WikiFre.gf ...` * **Purpose:** This produces the multi-lingual `semantik_architect.pgf`. * **Why:** GF cannot merge PGF files later. All languages must be present in the final Link command to be included in the binary. --- ## 6. How to Run the Pipeline ### Step 1: Audit the System (Update the Matrix) Run this whenever you add new files or change grammar code. ```bash # From project root python tools/everything_matrix/build_index.py ``` * **Output:** `data/indices/everything_matrix.json` ### Step 2: Build the Engine Run this to compile the PGF binary with AI assistance enabled. ```bash # Go to gf directory cd gf python builder/orchestrator.py ``` * **Output:** `gf/semantik_architect.pgf` * *Note:* Ensure `GOOGLE_API_KEY` is set in `.env` if you want the Architect Agent to fix errors. ### Step 3: Verify the Binary Check which languages actually made it into the binary. ```bash # Quick Python one-liner python3 -c "import pgf; print(pgf.readPGF('semantik_architect.pgf').languages.keys())" ``` * **Expected:** `['WikiEng', 'WikiFre', 'WikiZul', ...]` --- ## 7. Troubleshooting ### "My language isn't in the binary!" 1. **Check the Matrix:** Open `data/indices/everything_matrix.json`. Does your language exist? * *No?* Run `build_index.py`. * *Yes?* Check `build_strategy`. If it is `SKIP`, check the audit logs. 2. **Check Build Logs:** Look at `gf/build_logs/{lang}.log`. * *Common Error:* `File not found`. This means `rgl_auditor` detected the folder, but the actual `.gf` file path was unresolved. ### "The Architect Agent is consuming too much quota." * **Cause:** Frequent build failures triggering the repair loop. * **Fix:** Check `gf/build_logs/`. If a language is fundamentally broken, set its score to 0 manually in the Matrix (or delete the file) to stop the Agent from trying to fix it. ### "Lexicon score is 0 but I added files." * **Cause:** Your JSON structure might be invalid. * **Fix:** The `lexicon_scanner.py` requires valid JSON. If `json.load()` fails, it skips the file. Check your syntax. ================================================================================================ FILE: docs/Technical-Reference/03-LEXICON_ARCHITECTURE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 71b52df65b3dff70ce2b6d95f1f9578ac373d149f51d405cf2c3b64956f47442 CONTENT_BYTES: 6501 ================================================================================================ Here is the updated **`docs/03-LEXICON_ARCHITECTURE.md`**. This version strictly enforces the **ISO 639-1 (2-letter)** directory standard (`en` vs `eng`) across the file tree, examples, and scanner logic, aligning with the v2.1 "Code-First" mandate. --- # 📚 Lexicon Architecture & Workflow **SemantiK Architect v2.1** ## 1. Core Philosophy: Usage-Based Sharding Managing vocabulary for 300+ languages is a massive data challenge. A monolithic dictionary file (e.g., `french_all.json`) is inefficient and hard to maintain. We adopt a **Usage-Based Sharding** strategy: 1. **Upstream Source:** **Wikidata** is the "Raw Material." We use it to trace lineage (QIDs) and fetch translations. 2. **Downstream Usage:** **Domain Shards.** We organize local data by *semantic topic* (People, Science, Core). This allows the engine to load only the vocabulary needed for a specific context (e.g., loading `science.json` only when generating a Physics biography). 3. **Strict Validation:** Every entry must strictly adhere to a JSON schema to ensure the Grammar Engine never crashes due to missing attributes (like `gender`). --- ## 2. Directory Structure The lexicon lives in `data/lexicon/` and is organized hierarchically by **ISO 639-1 (2-letter)** language code. ```text data/ ├── lexicon/ │ ├── schema.json # Master Validation Schema (Draft-07) │ ├── en/ # English Namespace (NOT 'eng') │ │ ├── core.json # "Skeleton" words (is, the, he, she) │ │ ├── people.json # Professions, Titles, Relations │ │ ├── science.json # Scientific terms (Physics, Nobel Prize) │ │ └── geography.json # Countries, Cities, Demonyms │ ├── fr/ # French Namespace (NOT 'fra') │ │ ├── core.json │ │ └── ... │ └── zu/ # Zulu Namespace (NOT 'zul') │ └── ... └── imports/ # Staging area for bulk CSVs ├── en_wide.csv └── ... ``` --- ## 3. The Semantic Domains We divide vocabulary into four primary domains. The **Everything Matrix** scanner (`lexicon_scanner.py`) audits these specific files to calculate the "Zone B" readiness score. ### A. `core.json` (The Skeleton) * **Content:** Functional words required to construct *any* sentence. * **Examples:** Copulas ("is", "was"), Pronouns ("he", "it"), Articles ("the", "a"), Conjunctions. * **Criticality:** **Extreme.** If this file is missing or empty, the language is marked as **Broken** (Score 0). ### B. `people.json` (The Biography) * **Content:** Terms needed for the `BioFrame`. * **Examples:** * **Professions:** "Physicist", "Writer", "King". * **Relations:** "Spouse", "Child", "Advisor". * **Titles:** "Dr.", "PhD", "Sir". * **Criticality:** High. Required for the primary use case (Biographies). ### C. `geography.json` (The World) * **Content:** Location entities and their derived forms. * **Examples:** * **Entity:** "France" (Noun). * **Adjective:** "French" (Adj). * **Demonym:** "Frenchman" (Noun). * **Criticality:** Medium. Required for `nationality` fields. ### D. `science.json` (The Domain) * **Content:** Specialized terminology. * **Examples:** "Radioactivity", "Planet", "Theory of Relativity". * **Criticality:** Low (Initial), High (Production). --- ## 4. The Data Schema Every entry in the JSON files must validate against `data/lexicon/schema.json`. ### Base Entry Object ```json "physicist": { "pos": "NOUN", // Part of Speech: NOUN, VERB, ADJ, PROPN "qid": "Q169470", // Wikidata ID (Provenance) "gender": "m", // Grammatical Gender (m, f, n, c) - REQUIRED for Romance/Slavic "forms": { // Explicit overrides for irregulars "pl": "physicists" } } ``` ### Complex Types **1. Nationalities (`geography.json`)** Requires linking the country, the adjective, and the person-noun. ```json "french": { "pos": "ADJ", "qid": "Q142", // ID for "France" "forms": { "m": "français", "f": "française", "mpl": "français", "fpl": "françaises" }, "demonym": { // Link to the Noun form "m": "Français", "f": "Française" } } ``` **2. Verbs (`core.json`)** Requires conjugation stems if the language is not handled by the RGL smart paradigms. ```json "write": { "pos": "VERB", "qid": "Q223683", "forms": { "inf": "write", "past": "wrote", "pp": "written", "pres3sg": "writes" } } ``` --- ## 5. Audit & Maturity Scoring (Zone B) The **Everything Matrix** uses `lexicon_scanner.py` to grade the vocabulary readiness of every language. This score determines if a language is "Data Ready." | Score | Rating | Requirements | | --- | --- | --- | | **0** | 🔴 **Empty** | No JSON files found. | | **3** | 🟡 **Stub** | Files exist but contain `< 10` words. | | **5** | 🟠 **Minimal** | `core.json` exists (> 10 words). Can generate "A is B". | | **8** | 🔵 **Functional** | `core` + `people` exist (> 50 words). Can generate BioFrames. | | **10** | 🟢 **Production** | All domains present (> 200 words) OR a `_wide.csv` import exists. | **Impact on Build:** * If `lex_seed < 3`: The build pipeline marks `data_ready: false`. * **Auto-Correction:** The **Lexicographer AI** is triggered to generate the missing `core.json`. --- ## 6. Workflows ### Scenario A: Adding a New Language (Bootstrapping) 1. **Create Directory:** `mkdir data/lexicon/zu` (Zulu). 2. **Generate Seed:** Run the Lexicographer agent (or manually create `core.json`). 3. **Audit:** Run `python tools/everything_matrix/build_index.py` to confirm the scanner detects the new files. ### Scenario B: Bulk Import from Wikidata 1. **Fetch:** Use `scripts/fetch_wikidata_labels.py --lang=fr --domain=people`. 2. **Save:** The script outputs `data/imports/fr_wide.csv`. 3. **Audit:** The scanner detects the `.csv` and awards a **Zone B Score of 10**. 4. **Runtime:** The engine lazy-loads the CSV into memory during the first request. ### Scenario C: Fixing a "Missing Word" Error If the API returns `422 Unprocessable Entity`: 1. **Check Log:** The error message will say `Missing key: 'spaceman' in domain 'people'`. 2. **Edit:** Open `data/lexicon/{lang_code}/people.json`. 3. **Add:** Insert the entry for "spaceman". 4. **Restart:** You do **not** need to rebuild the PGF. Just restart the Worker/API to flush the JSON cache. ================================================================================================ FILE: docs/Technical-Reference/04-API_REFERENCE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 202b1b01a3d8359f8fe08add08f173ade29f2397c34c32988633d50878953b1a CONTENT_BYTES: 15891 ================================================================================================ # 🔌 API Reference & Semantic Frames **SemantiK Architect — Canonical Public HTTP Contract** Status: normative for the public HTTP generation surface Applies to: `/api/v1/generate/{lang_code}` and closely related public utility endpoints Related contracts: - `docs/contracts/public_generation_response_contract.md` - `docs/contracts/construction_runtime_contract.md` - `docs/contracts/debug_info_contract.md` - `docs/contracts/public_vs_runtime_vs_frontend_boundaries.md` - `docs/architecture/multilingual_runtime_target.md` - `docs/architecture/EN_FR_FINAL_PARALLEL_LOCKDOWN.md` --- ## 1. Overview SemantiK Architect exposes its text-generation runtime through a REST API. The canonical backend API is served under: - **Base path:** `/api/v1` - **Local dev default:** `http://localhost:8000/api/v1` - **Encoding:** UTF-8 - **Transport:** JSON over HTTP This reference defines the **canonical public HTTP contract** for generation and related utility endpoints. ### Canonical generation model The primary generation route is: **`POST /api/v1/generate/{lang_code}`** This route accepts a JSON object, normalizes it into the canonical internal semantic/frame shape, executes the planner-first runtime, and returns a structured JSON success response. ### Architectural notes - The backend is canonically mounted at `/api/v1/...`. - Health is intentionally available at both: - `/health/live`, `/health/ready` - `/api/v1/health/live`, `/api/v1/health/ready` - The canonical runtime path for successful generation is **planner-first**. - The canonical public success response is a structured JSON envelope centered on: - `text` - `lang_code` - `construction_id` - `renderer_backend` - `fallback_used` - `tokens` - `debug_info` - `generation_time_ms` - Older clients that depend primarily on `surface_text` / `meta` are not aligned with the canonical public contract. - Legacy input aliases may still be accepted at the request boundary, but they do not change the canonical public response shape. --- ## 2. Authentication This reference documents the API surface and payload/response contracts only. Authentication and authorization may be deployment-specific, for example: - reverse proxy enforcement, - API gateway enforcement, - protected admin/tooling routes, - environment-specific access policies. Do not assume that generation requires a built-in `X-API-Key` unless your deployment explicitly adds that requirement. --- ## 3. Primary Endpoint ## Generate Text **`POST /api/v1/generate/{lang_code}`** Generate natural-language text from a semantic payload. ### Path Parameters | Parameter | Type | Required | Description | | --- | --- | --- | --- | | `lang_code` | `string` | Yes | Authoritative language code for the request. The router normalizes supported variants. | ### Request Headers | Header | Value | Required | Description | | --- | --- | --- | --- | | `Content-Type` | `application/json` | Yes | Request body must be a JSON object. | | `Accept` | `application/json` | Recommended | The canonical public response contract is JSON. | ### Request Body The request body must be a **single JSON object**. Rules enforced by the request boundary: - the URL path language is authoritative; - if both URL language and payload language are provided, they must match after normalization; - if no path language is provided by some internal caller or alternate mounting path, the payload must carry a recognized language field; - transport-level language fields are normalized at the API boundary before semantic/frame parsing; - request compatibility aliases are a boundary concern only and do not redefine the internal runtime contract. Recognized payload language aliases may include: - `lang` - `language` - `lang_code` - `inputs.language` - `inputs.lang` - `inputs.lang_code` --- ## 4. Supported Input Modes The public request boundary supports multiple input styles that converge to one internal semantic/frame model. The stable public rule is: **JSON object in, canonical JSON success envelope out.** ### A. Bio / person payloads The following frame types are treated as bio/person-compatible inputs and normalized through the same compatibility boundary: - `bio` - `biography` - `entity.person` - `entity_person` - `person` - `entity.person.v1` - `entity.person.v2` ### Canonical bio example ```json { "frame_type": "bio", "name": "Alan Turing", "profession": "mathematician", "nationality": "British", "gender": "m" } ```` ### Compatibility bio example ```json { "frame_type": "entity.person.v2", "subject": { "name": "Alan Turing", "profession": "mathematician", "nationality": "British" } } ``` ### Common bio fields | Field | Type | Required | Notes | | ------------- | ---------------- | ------------ | ---------------------------------------------------- | | `frame_type` | `string` | Yes | Prefer `bio` for new clients. | | `name` | `string` | Usually yes | Common top-level compatibility field. | | `profession` | `string` | Commonly yes | May be normalized through lexical resolution. | | `nationality` | `string` | No | Optional. | | `gender` | `string \| null` | No | Optional compatibility field. | | `subject` | `object` | Sometimes | Used by some newer or compatibility person payloads. | ### B. Generic semantic frame payloads Non-bio semantic payloads are also accepted when they match supported internal frame/domain semantics. Example: ```json { "frame_type": "event", "subject": "Marie Curie", "event_type": "award", "date": "1903" } ``` ### C. Ninai / function-style payloads The request boundary may also accept Ninai-style or function-oriented payloads through the Ninai adapter. Example: ```json { "function": "ninai.constructors.Statement", "args": [ { "function": "ninai.types.Bio" }, { "function": "ninai.constructors.List", "args": ["physicist"] }, { "function": "ninai.constructors.Entity", "args": ["Q7186"] } ] } ``` This is a parsing/adapter concern only. The stable public generation contract remains the same. --- ## 5. Semantic Frame Rules Semantic frames are the canonical public input abstraction. ### Core rule Clients send semantic intent, not renderer-specific instructions. That means: * clients describe the meaning to be generated; * the planner/runtime selects the construction; * lexical resolution binds language-appropriate material; * the realizer/backend produces the final surface text. ### Public input boundary rules Public clients must not rely on: * GF-specific ASTs as the public contract, * backend-specific surface templates as the public contract, * direct renderer selection as a semantic requirement, * debug-only fields to carry required meaning. ### Practical guidance For new API clients: * prefer stable frame-style JSON objects; * prefer `frame_type: "bio"` for person/bio generation; * treat compatibility aliases as tolerated inputs, not as the long-term design center. --- ## 6. Success Response The canonical public success response is a JSON object with this top-level shape: | Field | Type | Required | Description | | -------------------- | ---------- | -------- | ----------------------------------------------------------- | | `text` | `string` | Yes | Final generated surface text. | | `lang_code` | `string` | Yes | Language code of the returned text. | | `construction_id` | `string` | Yes | Explicit construction identifier for the returned result. | | `renderer_backend` | `string` | Yes | Backend that produced the final text. | | `fallback_used` | `boolean` | Yes | Whether fallback was used in producing the returned result. | | `tokens` | `string[]` | Yes | Tokenized representation of the returned text. | | `debug_info` | `object` | Yes | Structured diagnostics object. | | `generation_time_ms` | `number` | Yes | Authoritative top-level generation time in milliseconds. | ### Canonical success response example ```json { "text": "Alan Turing is a British mathematician.", "lang_code": "en", "construction_id": "copula_equative_classification", "renderer_backend": "family", "fallback_used": false, "tokens": ["Alan", "Turing", "is", "a", "British", "mathematician."], "debug_info": { "runtime_path": "planner_first", "construction_id": "copula_equative_classification", "renderer_backend": "family", "lang_code": "en", "slot_keys": ["subject", "profession", "nationality"], "fallback_used": false, "selected_backend": "family", "attempted_backends": ["family"] }, "generation_time_ms": 12.5 } ``` ### Response rules * `text` is the authoritative public text field. * `lang_code` identifies the returned surface language. * `construction_id` is explicit on the canonical nominal path. * `renderer_backend` is explicit on the canonical nominal path. * `fallback_used` is explicit. * `tokens` correspond to the final text. * `generation_time_ms` is top-level and authoritative. * `debug_info` must not contradict top-level fields. * top-level nominal facts must not exist **only** inside `debug_info`. ### Parity rules When both top-level fields and `debug_info` carry the same fact, they must agree: * `lang_code == debug_info.lang_code` * `fallback_used == debug_info.fallback_used` * `renderer_backend == debug_info.renderer_backend` when both are present * `construction_id == debug_info.construction_id` when both are present --- ## 7. Diagnostics (`debug_info`) `debug_info` is the canonical structured diagnostics object carried in successful generation responses. ### Minimum stable shared diagnostics Canonical planner-first results preserve these stable shared keys when available: * `construction_id` * `renderer_backend` * `lang_code` * `slot_keys` * `fallback_used` * `runtime_path` ### Common recommended diagnostics Depending on runtime/backend availability, `debug_info` may also include: * `selected_backend` * `requested_backend` * `attempted_backends` * `dispatch_policy` * `fallback_reason` * `resolved_language` * `concrete_name` * `family` * `template_id` * `template_used` * `ast` * `lexical_resolution` * `backend_trace` * `warnings` * `errors` * `timings_ms` ### Diagnostics rules * `debug_info` is diagnostics only. * It is not a replacement for top-level public response fields. * It must be a JSON object. * It must be machine-readable first. * It must not contain secrets, credentials, or raw sensitive payload dumps. * Public serializers preserve diagnostics; they do not invent missing nominal planner-first facts. --- ## 8. Health Endpoints ### Live **`GET /health/live`** **`GET /api/v1/health/live`** Used for basic liveness checks. ### Ready **`GET /health/ready`** **`GET /api/v1/health/ready`** Used for readiness checks. Typical readiness response: ```json { "broker": "up", "storage": "up", "engine": "up" } ``` --- ## 9. Other Mounted Public Endpoints The app also mounts additional public API areas under `/api/v1`, including: * `/api/v1/languages` * `/api/v1/entities` * `/api/v1/frames` * `/api/v1/generate/{lang_code}` It may also mount protected, admin, or developer-oriented areas, including: * management endpoints under `/api/v1/...` * tools under `/api/v1/tools/...` This document focuses on the canonical generation contract. --- ## 10. Error Handling The generation route expects a JSON object and may reject invalid requests before generation starts. ### Common error situations | Status | Condition | | ------------- | ---------------------------------------------------------------------------- | | `400` / `422` | Invalid JSON object, invalid payload shape, or validation/parsing failure | | `400` | URL language and payload language do not match after normalization | | `400` | Missing language when no authoritative path language is provided | | `5xx` | Internal planner, lexical-resolution, realization, or infrastructure failure | ### Error-handling notes * Validation/parsing failures may occur before generation begins. * Runtime failures must not be smuggled through a success response. * Older error tables tied to exporter-specific behavior or obsolete transport assumptions must not be treated as authoritative for `POST /api/v1/generate/{lang_code}` unless they are explicitly reintroduced and implemented on this route. --- ## 11. Integration Guide (Python Client) ```python import requests API_BASE = "http://localhost:8000/api/v1" def generate_text(frame: dict, lang_code: str = "en") -> dict: url = f"{API_BASE}/generate/{lang_code}" response = requests.post( url, json=frame, headers={ "Content-Type": "application/json", "Accept": "application/json", }, timeout=30, ) response.raise_for_status() return response.json() result = generate_text( { "frame_type": "bio", "name": "Alan Turing", "profession": "mathematician", "nationality": "British" }, lang_code="en", ) print(result["text"]) print(result["lang_code"]) print(result["construction_id"]) ``` --- ## 12. Compatibility Notes ### Input compatibility Compatibility aliases may still be accepted at the request boundary, especially for bio/person-style payloads. Examples include: * alternate `frame_type` values for person/bio inputs, * nested `subject` forms, * Ninai-style parsing adapters, * transport-level language aliases. ### Output compatibility The canonical public response is the structured JSON envelope documented above. Clients should not depend on older assumptions such as: * top-level `surface_text` as the main success field, * top-level `meta` as the primary contract carrier, * text/plain as the canonical response contract for this route, * backend-specific internal fields as the public success contract. Compatibility handling may exist in code for migration-safe readers, but it does not redefine the public contract. --- ## 13. Deprecated Assumptions The following should be considered outdated for the canonical public generation route unless explicitly reintroduced and implemented: * response bodies centered only on `surface_text` / `meta` * `Accept: text/plain` as the primary documented contract * `Accept: text/x-conllu` as the documented contract for this route * `style` query parameter as part of the stable generation API contract * `X-Session-ID` as part of the stable generation API contract * legacy direct-frame execution as the canonical public runtime model --- ## 14. Summary The authoritative public generation contract is: * **Route:** `POST /api/v1/generate/{lang_code}` * **Input:** one JSON object carrying semantic/frame intent * **Runtime:** planner-first nominal generation * **Output:** one canonical JSON success envelope centered on: * `text` * `lang_code` * `construction_id` * `renderer_backend` * `fallback_used` * `tokens` * `debug_info` * `generation_time_ms` * **Diagnostics:** structured, machine-readable, and non-authoritative relative to top-level response fields * **Compatibility:** accepted at the request boundary where needed, but not allowed to redefine the canonical public contract ================================================================================================ FILE: docs/Technical-Reference/05-AI_SERVICES.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 4ecab865e8ab87bfe5087abadfebbac7cfaa9b5adba58b787a09c169fbb0133a CONTENT_BYTES: 7271 ================================================================================================ # 🧠 AI Services & Autonomous Agents **SemantiK Architect v2.5** ## 1. Overview The SemantiK Architect uses a **Hybrid Intelligence** model. * **The Core:** Deterministic, Rule-Based Engine (GF/Python). Guarantees grammatical correctness and verifiable output. * **The Edge:** Probabilistic AI Agents (LLMs). Handles high-entropy tasks like "guessing" a word's gender, grading translation naturalness, or serving as a **copilot** for grammar creation. * **The Paradigm Shift:** We have moved away from fully autonomous runtime generation in favor of a Human-in-the-Loop (HITL) model. This ensures determinism during builds and prevents API cost drains caused by LLM hallucinations. These agents are encapsulated in the `ai_services/` package and interact via the developer tools dashboard (`/tools`) or the CLI. --- ## 2. The Four Personas The architecture delegates responsibilities to four distinct AI agents. | Agent | Persona | Role | Trigger Event | | --- | --- | --- | --- | | **The Lexicographer** | *Data Generator* | Bootstraps dictionaries (`core.json`, `people.json`) for new languages. | `build_index.py` detects `lex_seed < 3`. | | **The Architect** | *The Copilot* | Generates concrete grammars for missing languages using a "GF Codex". | Invoked manually via `/tools` interface (`ai_refiner`) or `manage.py generate`. | | **The Surgeon** | *Code Fixer* | Suggests patches for broken `.gf` source files based on compiler error logs. | Manual diagnostic workflows (Removed from automatic build pipeline). | | **The Judge** | *QA Expert* | Grades the naturalness of the generated text against "Gold Standard" reference sentences. | Scheduled CI/CD or `test_quality.py`. | --- ## 3. Directory Structure & Configuration All AI logic is centralized to ensure consistent API handling and rate limiting. ```text ai_services/ ├── __init__.py # Exports the agents ├── client.py # Central Gemini Client (Auth, Rate Limiting, Backoff) ├── prompts.py # Frozen System Prompts and GF Codex (Source of Truth) ├── lexicographer.py # Logic for seeding dictionaries ├── architect.py # Logic for Generative Grammar Creation (HITL) ├── surgeon.py # Logic for Self-Healing suggestions └── judge.py # Logic for QA & Auto-Ticketing ``` **Environment Variables** The client requires the following `env` variables (configured in `app/shared/config.py`): * `GOOGLE_API_KEY`: Your Gemini API Key. * `AI_MODEL_NAME`: Defaults to `gemini-1.5-pro` (Reasoning optimized). * `GITHUB_TOKEN`: Required for The Judge to open issues. * `ARCHITECT_ENABLE_AI_TOOLS`: Must be set to `1` to allow execution of AI-gated tools via the backend router. --- ## 4. Agent Definitions ### A. The Lexicographer (`lexicographer.py`) **Goal:** Ensure no language has an empty dictionary. **Workflow:** 1. **Trigger:** The `lexicon_scanner.py` reports a language (e.g., `zul`) has `seed_score: 0`. 2. **Prompting:** The agent receives a list of core concepts ("is", "the", "person", "water"). 3. **Generation:** It asks the LLM to generate the JSON entries, including morphological features (e.g., *Zulu noun class prefixes*). 4. **Output:** Writes `data/lexicon/zul/core.json`. ### B. The Architect (`architect.py`) [UPDATED v2.5] **Goal:** Serve as a human-guided copilot to draft new GF grammars, drastically reducing manual boilerplate. **Workflow:** 1. **Trigger:** An operator identifies a missing language and launches the interactive `ai_refiner` tool from the developer dashboard (`/tools`). 2. **Prompting:** The tool sends the "GF Codex" and the typological order (SVO, SOV) of the target language to the LLM API. 3. **The GF Codex:** The AI relies on strict context injection including: * **Anti-Crash Rules:** E.g., The "Inlining Rule" (banning `let` inside `lin`) and the "Symbolic Rule" (banning `mkPN` in favor of `symb` for raw strings). * **Strict Skeletons:** Mandatory use of `Predicate = VP ;`. * **Few-Shot Examples:** Perfect templates of validated GF grammars. 4. **Human Validation:** The operator receives the draft, reviews it, and compiles it via `tools/language_health.py --mode compile`. 5. **Output:** Once successfully compiled, the grammar is permanently saved to `gf/contrib/{lang}/Wiki{Lang}.gf` (Manual Overrides). ### C. The Surgeon (`surgeon.py`) [UPDATED v2.5] **Goal:** Diagnostic assistance for broken grammar code. **Workflow Changes:** * The automated AI fallback loop that was previously triggered when compilation failed has been explicitly removed from `builder/orchestrator.py`. * This removes the "Fire and Pray" instability and endless retry loops during standard build processes. * The Surgeon is now invoked strictly as an interactive tool by developers trying to patch complex syntax errors. ### D. The Judge (`judge.py`) **Goal:** Quality Assurance beyond "It compiles." **Workflow:** 1. **Trigger:** The `test_quality.py` script runs a regression test. 2. **Reference:** It loads the **Gold Standard** dataset from `data/tests/gold_standard.json` (migrated from Udiron). 3. **Comparison:** It compares the SKA generation against the ground truth. 4. **Action:** * **Pass:** `similarity > 0.8`. * **Fail:** `similarity < 0.8`. 5. **Whistleblowing:** If it fails with high confidence, it calls the GitHub API to open an issue automatically. **Issue Template:** > **Title:** `[QA] Poor Quality: {Lang} - {Frame}` > **Body:** "Expected 'Shaka is a warrior', got 'Me Shaka warrior'. Confidence: 95%." --- ## 5. Integration Hooks (The Deterministic Pipeline) Because we have shifted to a Human-in-the-Loop model, the automated `if not compile_success:` AI hooks have been removed from the build orchestrator. **The New Pipeline Behavior:** 1. **Deterministic Build:** During the regular pipeline (`build_300.py` or `orchestrator.py`), the orchestrator looks for verified files in `contrib/` and links them directly into the `semantik_architect.pgf` binary. 2. **Graceful Degradation:** If a language is broken or missing, the orchestrator simply skips it (`SKIP`) and logs an alert that human intervention is required. 3. **Zero API Calls:** No LLM API calls are made during the build sequence, ensuring 100% determinism and eliminating mid-build latency. **CI/CD Integration (QA):** The **Judge** continues to be triggered safely via the test runner: ```bash # Runs the full regression suite using the Judge Agent python -m pytest tests/integration/test_quality.py --use-judge ``` --- ## 6. Rate Limiting & Cost Control The `client.py` module implements a robust **Exponential Backoff** strategy to handle API quotas during interactive sessions. * **Retries:** Max 3 attempts per request. * **Backoff:** 2s -> 4s -> 8s delay between retries. * **Circuit Breaker:** If 5 consecutive requests fail, the AI service disables itself to prevent credit drain. * **Overall Impact:** Moving the Architect to the HITL model drastically reduces overall API costs, as calls are only made once per language rather than on every automated build. --- ## 7. Future Roadmap * **Learned Micro-Planning:** Using the LLM to rewrite frame parameters (e.g., synonyms) for stylistic variation *before* rendering. ================================================================================================ FILE: docs/Technical-Reference/06-ADDING_A_LANGUAGE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 50cbc21623f059cefde9b27b00f204e74b3ecf8bc88da3e9cadeec6c85e8d3cf CONTENT_BYTES: 6347 ================================================================================================ # 🌍 Adding a New Language **SemantiK Architect v2.5** This guide documents the standard workflow for adding support for a new language (e.g., `pt` for Portuguese or `ha` for Hausa). The process is: 1. register the language (so it shows up in the Everything Matrix), 2. seed lexicon data, 3. ensure topology config exists (Tier 3), 4. build and confirm the PGF contains the language, 5. add QA coverage. --- ## 🏗️ Phase 1: Registration (The Matrix) Before you write any grammar code, register the language so the build system knows it exists. ### Step 1: Identify the Tier * **Tier 1 (High Road):** Supported by GF RGL (Resource Grammar Library) and can be bootstrapped/aligned. * **Tier 3 (Safe Mode / Factory):** Not in RGL; uses the deterministic Grammar Factory SAFE_MODE output. ### Step 2: Ensure ISO → Wiki suffix mapping is correct The system uses ISO language codes (usually ISO-639-1 like `pt`, `de`, `ha`) but GF module names use a **Wiki suffix** (e.g., `WikiPor`, `WikiGer`, `WikiEng`). This mapping is authoritative in: * `data/config/iso_to_wiki.json` (preferred location) Add/update the entry for your ISO code: ```json { "pt": { "wiki": "Por" }, "de": { "wiki": "Ger" }, "ha": { "wiki": "Ha" } } ``` Notes: * For **Tier 1** languages, this mapping is usually **required** (RGL suffixes often do *not* match `code.title()`). * For **Tier 3** languages, if you omit it, the system falls back to `TitleCase` (`ha` → `Ha` → `WikiHa`). ### Step 3: Register Tier 3 languages in the Factory Wishlist (only if Tier 3) If the language is not in RGL, register it so the system can generate SAFE_MODE grammar: 1. Open `utils/grammar_factory.py`. 2. Add the ISO 2-letter code: ```python "ha": {"name": "Hausa", "order": "SVO", "family": "Chadic"} ``` The `order` field (`SVO`, `SOV`, `VSO`, …) drives Weighted Topology behavior. ### Step 4: Rebuild the Everything Matrix Run the indexer so the language appears in the matrix: ```bash python tools/everything_matrix/build_index.py --langs ha ``` Verify in `data/indices/everything_matrix.json`: * `verdict.build_strategy` should be: * Tier 1: `"HIGH_ROAD"` * Tier 3: `"SAFE_MODE"` --- ## 📚 Phase 2: The Lexicon (The Data) A grammar without words is useless. Populate Zone B lexicon files under the 2-letter directory: ### Manual creation **File (mandatory):** `data/lexicon/ha/core.json` ```json { "verb_be": { "pos": "VERB", "lemma": "ne", "forms": { "pres_3sg": "ne", "past_3sg": "ne" } } } ``` **File (common):** `data/lexicon/ha/people.json` ```json { "physicist": { "pos": "NOUN", "qid": "Q169470", "forms": { "sg": "masanin kimiyyar", "pl": "masana kimiyya" } } } ``` ### Optional: AI-assisted seeding If enabled in your environment, use the lexicon seeding utility to fill gaps: ```bash python utils/seed_lexicon_ai.py --langs ha --domains core people --apply ``` --- ## ⚙️ Phase 3: Configuration (Topology) Tier 3 SAFE_MODE grammars use **Weighted Topology**. 1. Open `data/config/topology_weights.json` 2. Ensure your chosen word order exists: ```json { "SVO": { "nsubj": -10, "root": 0, "obj": 10 }, "SOV": { "nsubj": -10, "obj": -5, "root": 0 } } ``` If the language uses a rare order, add a new entry. --- ## 🚀 Phase 4: Build & Deploy ### Step 1: Build (preferred entrypoint) This is the recommended “do the right thing” build entry: ```bash python manage.py build --langs ha --align ``` * `--langs ha` scopes work to the language you’re adding. * `--align` performs Tier 1 bootstrap/alignment where applicable, and ensures build inputs are in place. ### Step 2: Orchestrator only (compile + link pipeline) If you only want the orchestrator step: ```bash python -m builder.orchestrator --strategy AUTO --langs ha ``` What happens: * **AUTO** uses `data/indices/everything_matrix.json` verdicts to choose `"HIGH_ROAD"` vs `"SAFE_MODE"` per language. * Build is **two-phase**: compile individual `.gf` → link into `gf/semantik_architect.pgf`. ### Step 3: Verify the binary contains the language ```bash python3 -c "import os, pgf; p=os.getenv('PGF_PATH','gf/semantik_architect.pgf'); g=pgf.readPGF(p); print(sorted(g.languages.keys()))" ``` Expected: an entry like `WikiHa` (or `WikiPor`, `WikiGer`, etc. depending on `iso_to_wiki.json`). --- ## 🧪 Phase 5: Quality Assurance (The Judge) ### Step 1: Add a Gold Standard test Open `data/tests/gold_standard.json` and add: ```json { "lang": "ha", "intent": { "frame_type": "bio", "name": "Shaka", "profession": "warrior" }, "expected": "Shaka jarumi ne." } ``` ### Step 2: Run regression ```bash python -m pytest tests/integration/test_quality.py --lang=ha ``` ### Step 3: Smoke test (API) ```bash curl -X POST "http://localhost:8000/api/v1/generate/ha" \ -H "Content-Type: application/json" \ -d '{ "frame_type": "bio", "name": "Shaka", "profession": "warrior" }' ``` --- ## 📝 Summary Checklist | Phase | Action | Verification | | --------------- | --------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | | **1. Register** | Add ISO→Wiki mapping in `data/config/iso_to_wiki.json` (and add to `grammar_factory.py` if Tier 3). | `everything_matrix.json` includes language; `build_strategy` correct. | | **2. Lexicon** | Create `data/lexicon//*.json` (or seed via `utils/seed_lexicon_ai.py`). | Lexicon loads; coverage improves. | | **3. Config** | Ensure `topology_weights.json` supports the language’s order (Tier 3). | N/A | | **4. Build** | `python manage.py build --langs xx --align` (or `python -m builder.orchestrator ...`). | PGF includes `Wiki` language key. | | **5. QA** | Add gold standard + run tests. | Tests pass; output acceptable. | ================================================================================================ FILE: docs/Technical-Reference/07-LINGUISTICS_REFERENCE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 4cb3a3f652a330d99858916213da60eaebcb254c18c0e4f3c9925dbbb8de4057 CONTENT_BYTES: 5676 ================================================================================================ # 🎓 Linguistics Reference & Theory **SemantiK Architect v2.0** ## 1. Purpose & Positioning This document explains **why** the system is built the way it is. It maps the engineering components (Engines, Matrices, JSON) to their corresponding concepts in linguistic theory. The SemantiK Architect v2.0 is designed to be: * **Engineered for Scale:** Capable of supporting 300+ languages through automation. * **Interoperable:** Native support for the **Ninai** protocol and **Universal Dependencies (UD)** standards. * **Theory-Aware:** Informed by research-grade formalisms to ensure it can handle the complexity of natural language. ### High-Level Analogy The system sits at the intersection of four traditions: 1. **Grammatical Framework (GF):** Separation of *Abstract Syntax* (Logic) and *Concrete Syntax* (Strings). 2. **Dependency Grammar:** We use "Weighted Topology" (Udiron) to linearize text based on head-dependent relationships. 3. **Frame Semantics:** We use **Ninai Constructors** as the atomic unit of meaning. 4. **Discourse Theory:** We explicitly model entity salience (Centering Theory) to handle pronouns. --- ## 2. Theoretical Foundations ### 2.1 The Ninai Protocol (Abstract Syntax) In v2.0, we adopt **Ninai** (Abstract Wikipedia's notation) as our primary representation of meaning. * **Concept:** Language-Independent Logic Form. * **Implementation:** `app/adapters/ninai.py`. * **Theory:** A Ninai Object Tree represents the *Deep Structure* of a sentence. It uses **Constructors** (e.g., `ninai.constructors.Statement`) to define relationships without committing to a specific word order or morphology. ### 2.2 Grammatical Framework (Concrete Syntax) We use GF as the low-level runtime engine to realize the Deep Structure into Surface Text. * **Role:** Handles the "Morphological Explosion" (e.g., Finnish noun cases). * **Hybridization:** * **Tier 1 (RGL):** Uses "Hand-Written Grammars" (Chomskyan / Generative). * **Tier 3 (Factory):** Uses "Topology Grammars" (Data-Driven). ### 2.3 Universal Dependencies (Evaluation) We bridge the gap between **Generative Grammar** (building trees) and **Dependency Grammar** (analyzing links). * **Theory:** "Construction-Time Tagging." * **Implementation:** `app/core/exporters/ud_mapping.py`. * **Logic:** Since we *build* the sentence, we know exactly which word is the Subject (`nsubj`) and which is the Object (`obj`). We map these intents to **CoNLL-U** tags dynamically, allowing our output to be evaluated against standard treebanks. --- ## 3. Tier 3 Theory: Weighted Topology (Udiron) For under-resourced languages, writing a full generative grammar is too slow. We adopt the **Weighted Topology** approach from the `Udiron` project. ### 3.1 The Linearization Problem How do you generate text for 300 languages when some are SVO (English), some SOV (Japanese), and some VSO (Irish)? ### 3.2 The Topological Solution We view a sentence not as a tree, but as a **Field of Slots** sorted by weight. * **The Mechanism:** We assign integer weights to dependency roles relative to the Root (Verb). * **Configuration:** `data/config/topology_weights.json`. **Example: Subject-Object-Verb (SOV)** * `Subject (nsubj)`: **-10** (Far Left) * `Object (obj)`: **-5** (Left) * `Verb (root)`: **0** (Center) **Result:** The engine simply sorts the constituents by weight: `[-10, -5, 0]` `Subject + Object + Verb`. This allows us to support any word order configuration purely through configuration, without changing code. --- ## 4. Discourse & Context (Centering Theory) In v2.0, we moved beyond single sentences to **Discourse Planning**. ### 4.1 The Problem * Sentence 1: "Marie Curie is a physicist." * Sentence 2: "Marie Curie was born in Poland." * *Critique:* Repetitive and unnatural. ### 4.2 The Solution: Entity Salience We implement a simplified version of **Centering Theory**. * **Backward-Looking Center ():** The entity currently "in focus" from the previous utterance. * **Implementation:** `SessionContext` in Redis. * **Rule:** If the **Subject** of the current sentence matches the **** of the session, we apply a **Pronominalization Transformation** (Swap Name Pronoun). --- ## 5. Design Trade-offs ### 5.1 Determinism vs. Variation (Micro-Planning) * **The Tension:** Rule-based systems are repetitive. AI systems are hallucination-prone. * **The v2.0 Compromise:** **Learned Micro-Planning**. * We use **AI (LLM)** to select the *lexical items* (Style). * We use **GF (Rules)** to assemble the *syntax* (Grammar). * *Result:* We can vary "died" vs "passed away" (Style) without risking "He passed away" for a female subject (Grammar handles gender). ### 5.2 NLG-First * **The Choice:** The system is strictly **Generation-First** (NLG). * **Implication:** We do not parse text. We render data. This eliminates the "Ambiguity Problem" common in translation systems because the input (Ninai JSON) is unambiguous by design. --- ## 6. Summary of Systems | Component | Linguistic Concept | Implementation | | --- | --- | --- | | **Ninai Adapter** | Deep Structure / Logical Form | Recursive JSON Parser | | **Lexicon** | Lexical Semantics | JSON Shards (`people.json`) | | **RGL (Tier 1)** | Generative Grammar | `.gf` Source Files | | **Factory (Tier 3)** | Topological Fields / Linearization | `topology_weights.json` | | **Discourse Planner** | Centering Theory / Coreference | Redis Session Store | | **UD Exporter** | Dependency Grammar | `ud_mapping.py` | This document confirms that the SemantiK Architect v2.0 is a **Hybrid Neuro-Symbolic System**, leveraging the best of formal linguistics (GF/UD) and modern engineering (Redis/AI). ================================================================================================ FILE: docs/Technical-Reference/08-DECISION_LOG.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 1e99ced208b6d132ca57702cdd225bacdbae7f4a9163da6c57f0b77b9befc12b CONTENT_BYTES: 5879 ================================================================================================ # 📜 Decision Log & Architecture Records **SemantiK Architect v2.0** This document records the **key architectural choices** behind the SemantiK Architect and the alternatives that were considered. It is meant to be a concise "Why we did it this way" reference for reviewers and future contributors. --- ## 1. Overall System Shape ### Decision: Hexagonal Architecture (Ports & Adapters) **The Context** NLG systems can be built as monoliths or pipelines. We needed a structure that could handle 300+ languages and integrate with external AI agents and standards (Ninai, UD) without tight coupling. **The Decision** We adopted a **Hexagonal Architecture**: 1. **Core Domain:** Pure Python logic (`BioFrame`, `GrammarEngine`). 2. **Input Ports:** `NinaiAdapter` (JSON Trees) and `API` (HTTP). 3. **Output Ports:** `UDMapping` (CoNLL-U) and `TextRenderer`. 4. **Adapters:** Redis (State), GF Runtime (C-Bindings), Gemini (AI). **Why we chose this** * **Interoperability:** We can swap the input format from our internal JSON to the **Ninai Protocol** without touching the core linguistic logic. * **Testing:** We can test the Core Domain without spinning up the C-runtime or Redis. --- ## 2. Input Protocol ### Decision: Adopting Ninai (Recursive Objects) vs. Flat JSON **The Context** v1.0 used a flat `BioFrame`. Abstract Wikipedia uses **Ninai**, a recursive LISP-like object structure (Constructors). **The Decision** We built the **Ninai Bridge (`app/adapters/ninai.py`)** to transform recursive Ninai objects into our flat internal frames, rather than rewriting the entire engine to work natively on trees. **Why we chose this** * **Standardization:** Allows SKA to function as a compliant renderer for the Abstract Wikipedia ecosystem. * **Stability:** Keeps our internal domain logic simple (flat) while supporting complex external inputs (trees). --- ## 3. Tier 3 Linearization (The Factory) ### Decision: Weighted Topology (Udiron) vs. Hardcoded Templates **The Context** In v1.0, the "Factory" generated hardcoded `SVO` string concatenation. This produced grammatically incorrect output for SOV languages (Japanese) or VSO (Irish). Writing custom code for each was unscalable. **The Decision** We adopted **Weighted Topology Sorting** (adapted from the **Udiron** project). We assign relative integer weights to dependency roles (e.g., `subj=-10`, `obj=-5`, `verb=0` for SOV) and sort them at runtime. **Why we chose this** * **Zero-Code Config:** We can support *any* word order (OVS, VOS, etc.) just by editing `topology_weights.json`. * **Simplicity:** The factory logic remains generic; only the weights change per language. --- ## 4. Evaluation Standard ### Decision: Construction-Time Tagging (UD) vs. Post-Hoc Parsing **The Context** To prove our output is "good," we need to evaluate it. Running a 3rd-party dependency parser on our output is slow and error-prone (parsing is guessing). **The Decision** We implemented **Construction-Time Tagging**. Since we *build* the sentence using specific functions (`mkCl`, `mkNP`), we know exactly what is a Subject and what is an Object. We map these intents directly to **Universal Dependencies (CoNLL-U)** tags. **Why we chose this** * **Accuracy:** 100% accurate tagging because it is based on the generator's intent, not a parser's guess. * **Speed:** Zero runtime overhead compared to loading a Neural Parser. --- ## 5. State Management ### Decision: Redis Session Store vs. Stateless Requests **The Context** Generating isolated sentences leads to repetition ("Marie Curie is X. Marie Curie is Y."). We needed to implement **Pronominalization** (using "She"). **The Decision** We introduced **Redis** to store a `SessionContext` (ID + History). The **Discourse Planner** checks this context to decide whether to render a Name or a Pronoun. **Why we chose this** * **Performance:** Redis is sub-millisecond, essential for a real-time NLG API. * **Decoupling:** The renderer doesn't need to know *why* it's rendering "She," just that the context dictates it. --- ## 6. The "Everything Matrix" (Data-Driven Build) ### Decision: Dynamic System Scanning vs. Static Config **The Context** Hardcoding `LANGS = ['eng', 'fra']` leads to configuration drift. **The Decision** We built the **Everything Matrix**, a dynamic registry populated by scanning the filesystem before every build. **Why we chose this** * **Truth:** The build system never lies. If the file isn't on disk, the Matrix marks it `BROKEN`. * **Automation:** Adding a language is as simple as adding the files; the scanner auto-registers it. --- ## 7. AI Services Integration ### Decision: "The Architect" & "The Judge" Agents **The Context** * **Problem A:** Writing grammar files for 300 languages is too much work. * **Problem B:** We can't manually verify quality for 300 languages. **The Decision** We integrated specialized AI Agents: * **The Architect:** Generates the `.gf` code for Tier 3 languages using the **Frozen System Prompt**. * **The Judge:** Validates output against **Gold Standard** data and auto-files GitHub issues. **Why we chose this** * **Scale:** AI acts as a force multiplier, writing code and checking quality faster than humans. * **Consistency:** The System Prompt ensures the AI writes deterministic GF code, not chatty markdown. --- ## 8. Summary The key design choices defining v2.0 are: 1. **Hexagonal Architecture:** For Ninai/UD interoperability. 2. **Weighted Topology:** For solving the "Word Order" problem without code. 3. **Redis Context:** For Discourse Planning (Pronouns). 4. **Hybrid Factory:** Combining RGL (Expert) and AI Architect (Automated). 5. **Construction-Time Tagging:** For reliable evaluation. Together, these choices create a system that is **scalable (300+ languages), interoperable (Standard Protocols), and autonomous (AI-Driven)**. ================================================================================================ FILE: docs/Technical-Reference/09-AI_CONTEXT_DUMP.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 3d8339a474ec44cce1369137180df4442e705d33ea49c845916bfef16a7f5a97 CONTENT_BYTES: 4517 ================================================================================================ # SYSTEM CONTEXT: SemantiK Architect v2.0 ("Omni-Upgrade") # ============================================================================== # INSTRUCTIONS FOR AI: # You are acting as the Lead Architect for the "SemantiK Architect" project. # This is a Hybrid Neuro-Symbolic NLG Engine combining: # 1. Grammatical Framework (GF) -> Deterministic Rule-Based Core. # 2. Ninai Protocol -> Recursive Abstract Syntax Input. # 3. AI Agents (LLMs) -> Autonomous Code Repair & Grammar Generation. # # REFERENCE THIS CONTEXT FOR ALL FUTURE RESPONSES. # ============================================================================== # 1. CORE ARCHITECTURE # ------------------------------------------------------------------------------ # STYLE: Hexagonal Architecture (Ports & Adapters). # INPUT PORTS: # - Semantic Frame (Internal Flat JSON). # - Ninai Protocol (External Recursive JSON Object Tree). # ENGINE: Hybrid Factory. # - Tier 1: RGL (Resource Grammar Library) -> High Quality (Expert). # - Tier 3: Weighted Topology (Udiron-based) -> Automated Linearization. # STATE: Redis-backed "Discourse Planner" (SessionContext) for Pronominalization. # OUTPUT PORTS: Text (String) and Universal Dependencies (CoNLL-U). # PIPELINE: Two-Phase Build (Verify -> Link) + Architect Agent Repair Loop. # 2. FILE SYSTEM & PATHS (Hybrid WSL) # ------------------------------------------------------------------------------ # ROOT (Windows): C:\MyCode\SemantiK_Architect\Semantik_architect # ROOT (Linux): /mnt/c/MyCode/SemantiK_Architect/Semantik_architect # DOCKER MOUNT: /app # KEY DIRECTORIES: # gf/ -> Build artifacts (semantik_architect.pgf) & Orchestrator. # gf-rgl/src/ -> Tier 1 Source (External submodule). # generated/src/ -> Tier 3 Source (AI/Factory generated). # data/indices/ -> "Everything Matrix" (System Registry). # data/lexicon/{iso}/ -> Vocabulary shards (core.json, people.json). # data/config/ -> Topology Weights (SVO/SOV definitions). # data/tests/ -> Gold Standard QA Data (migrated from Udiron). # app/adapters/ -> Ninai Bridge, API, Redis Bus. # ai_services/ -> Autonomous Agents (Architect, Judge, Surgeon). # 3. CRITICAL FILES & SCRIPTS # ------------------------------------------------------------------------------ # BUILDER: builder/orchestrator.py # (Runs Verify->Link loop. Triggers 'Architect Agent' on failure). # FACTORY: utils/grammar_factory.py # (Implements Weighted Topology sorting for Tier 3). # AUDITOR: tools/everything_matrix/build_index.py # (Scans FS, updates everything_matrix.json). # NINAI: app/adapters/ninai.py # (Recursive JSON parser for Ninai Object Model). # MAPPING: app/core/exporters/ud_mapping.py # (Frozen Dictionary mapping RGL functions to UD Tags). # QA: ai_services/judge.py # (Validates output against 'gold_standard.json'). # 4. DATA CONTRACTS (JSON SCHEMAS) # ------------------------------------------------------------------------------ # NINAI INPUT (Recursive): # { "function": "ninai.constructors.Statement", "args": [...] } # # SEMANTIC FRAME (Internal): # { "frame_type": "bio", "name": "X", "profession": "Y", "context_id": "UUID" } # # SESSION CONTEXT (Redis): # { "session_id": "UUID", "current_focus": { "qid": "Q1", "gender": "f" } } # # TOPOLOGY WEIGHTS (Config): # { "SOV": { "nsubj": -10, "obj": -5, "root": 0 } } # 5. ENVIRONMENT VARIABLES (.env) # ------------------------------------------------------------------------------ # APP_ENV=development # REDIS_URL=redis://redis:6379/0 # Replaces old REDIS_HOST # SESSION_TTL_SEC=600 # Context duration # GITHUB_TOKEN=ghp_... # For Judge Agent Auto-Ticketing # GOOGLE_API_KEY=AIza... # For Architect/Surgeon Agents # REPO_URL=https://github.com/... # Target for Issue Creation # 6. KNOWN CONSTRAINTS & RULES # ------------------------------------------------------------------------------ # 1. NO WINDOWS RUNTIME: Backend MUST run in WSL/Linux (libpgf dependency). # 2. VARIABLE LEDGER: Use strict variable names from 'docs/14-VAR_FIX_LEDGER.md'. # 3. FROZEN PROMPTS: Do not invent LLM prompts. Use 'ai_services/prompts.py'. # 4. NINAI PROTOCOL: Input is Object-Based (JSON), NOT Lisp-String based. # 5. NO HARDCODED GRAMMAR: Use 'topology_weights.json' for word order. ================================================================================================ FILE: docs/Technical-Reference/10-GLOSSARY.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: ab57013ca79adad73b4555113cd157eada976ea7051d330b3b71ae57dfb50c99 CONTENT_BYTES: 3925 ================================================================================================ # 📖 Project Glossary & Terminology **SemantiK Architect** This document defines the specialized terminology used across the project. It bridges the gap between **Software Engineering** concepts and **Computational Linguistics** concepts. --- ## 🏛️ System Architecture Terms ### **Everything Matrix** * **Definition:** The dynamic registry (`everything_matrix.json`) that tracks the maturity and build status of every language in the system. * **Context:** It replaces static configuration. The system scans the filesystem to populate this matrix. * **Code:** `tools/everything_matrix/` ### **Hexagonal Architecture** * **Definition:** A design pattern (Ports & Adapters) that isolates the core domain logic from external tools like databases or APIs. * **Context:** We use this to ensure the `BioFrame` logic doesn't care if the grammar is stored in S3 or on the local disk. * **Code:** `app/core/` (Domain), `app/adapters/` (Infrastructure). ### **Hybrid Factory** * **Definition:** The strategy of combining expert-written grammars (Tier 1) with auto-generated simplified grammars (Tier 3) to achieve 100% language coverage. * **Context:** Used by the `builder/orchestrator.py` to decide which source files to include. ### **Two-Phase Build** * **Definition:** The compilation strategy used to solve the "Last Man Standing" bug. 1. **Verify:** Compile individual languages to temporary object files (`.gfo`). 2. **Link:** Merge all valid objects into a single binary (`.pgf`). --- ## 🗣️ Linguistic & GF Terms ### **Abstract Syntax** * **Definition:** The logical "skeleton" of a grammar. It defines *what* can be said (e.g., "A Sentence consists of a Subject and a Predicate") without defining *how*. * **Context:** Defined in `gf/semantik_architect.gf`. It is the language-independent interface. ### **Concrete Syntax** * **Definition:** The language-specific implementation of the Abstract Syntax. It defines *how* to say it (e.g., "In French, the adjective comes after the noun"). * **Context:** Defined in `WikiFre.gf`, `WikiEng.gf`. ### **Linearization** * **Definition:** The process of turning a tree structure (Abstract Syntax) into a flat string of text (Concrete Syntax). * **Context:** This is what the `pgf` C-runtime does when the API is called. ### **Morphology** * **Definition:** The study of the internal structure of words (inflection). * **Context:** Handling how "run" becomes "ran" or how "gato" (cat) becomes "gatos" (cats). * **Code:** Handled by the **RGL** (Tier 1) or simple string concatenation (Tier 3). ### **PGF (Portable Grammar Format)** * **Definition:** The compiled binary format of a GF grammar. It is to GF what `.class` is to Java. * **Context:** The file `gf/semantik_architect.pgf` is the final artifact loaded by the API. ### **RGL (Resource Grammar Library)** * **Definition:** The standard open-source library for GF that implements the morphology and syntax of ~40 languages. * **Context:** We use this as our "High Road" (Tier 1) source. --- ## 💾 Data & Logic Terms ### **Domain Sharding** * **Definition:** Splitting the vocabulary into small, topic-specific files (`science.json`, `people.json`) instead of one giant dictionary. * **Context:** Optimizes memory usage by only loading relevant terms. ### **Semantic Frame** * **Definition:** A JSON object representing an abstract intent (e.g., `BioFrame`, `EventFrame`). * **Context:** This is the input to the API. It is language-agnostic. ### **Saga Pattern** * **Definition:** A way to manage long-running transactions in a distributed system (e.g., "Start Build" -> "Wait" -> "Update Worker"). * **Context:** Used implicitly by our Async Worker to handle grammar hot-reloading. ### **Tier 1 / Tier 3** * **Tier 1:** Mature language backed by the RGL. High quality. * **Tier 3:** "Pidgin" language backed by the Factory. Lower quality (SVO only), but guarantees availability. ================================================================================================ FILE: docs/Technical-Reference/11-CONTRIBUTING.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: cfb9c5e11619ebfca6dd4cce2bffab6b67ccf5e3982b410ee0378a773e4f3459 CONTENT_BYTES: 4140 ================================================================================================ # 🤝 Contributing to SemantiK Architect v2.0 Thank you for your interest in contributing! This project is a complex **Hybrid Neuro-Symbolic Architecture** (Python/WSL + C/Linux + AI Agents), so we enforce strict guidelines to prevent "Works on my machine" issues. ## 1. The Golden Rules 1. **Never Edit Generated Files:** Do not manually edit files in `generated/src/`. These are overwritten by the **Architect Agent** and `builder/orchestrator.py`. 2. **Respect the Matrix:** Do not hardcode language lists in Python. If you add a language, register it by creating the file structure and running `tools/everything_matrix/build_index.py`. 3. **Linux Runtime Only:** The backend depends on `libpgf` (C-library). Do not try to run `uvicorn` or `worker.py` directly on Windows. Use WSL 2 or Docker. 4. **No Chatty AI:** If you modify the **Architect Agent**, you must strictly adhere to the **Frozen System Prompt** (`ai_services/prompts.py`) to ensure it outputs raw code, not Markdown. --- ## 2. Development Workflow ### Adding a New Language (The v2.0 Way) 1. **Tier 1 (High Quality):** Ensure the language exists in `gf-rgl/src`. 2. **Tier 3 (Factory):** * Add the language code and `order` (e.g., SVO, SOV) to `utils/grammar_factory.py`. * Verify the **Topology Weights** exist in `data/config/topology_weights.json`. 3. **Lexicon:** * Run `python -m ai_services.lexicographer --lang={code}` to bootstrap `core.json`. 4. **Audit:** Run `python tools/everything_matrix/build_index.py`. 5. **Build:** Run `python builder/orchestrator.py`. (The **Architect Agent** will wake up and write the grammar for you). ### Reporting Bugs * **Label:** Use `[Engine]` for GF/PGF issues, `[API]` for FastAPI, `[Matrix]` for scanners, and `[AI]` for agent hallucinations. * **Context:** Always include the `trace_id` and the **Session ID** if the bug involves Pronominalization/Context. --- ## 3. Coding Standards ### Python (Backend) * **Style:** We use `black` for formatting. * **Typing:** Strict type hints (`mypy`) are required for all `app/core` logic. * **Hexagonal:** Domain logic (`app/core`) must **never** import from `app/adapters`. * **Variables:** Use the **Frozen Ledger** (`docs/14-VAR_FIX_LEDGER.md`) for all shared constants. ### GF (Grammar) * **Naming:** Concrete grammars must be named `Wiki{Lang}.gf` (e.g., `WikiZul.gf`). * **Paradigm:** Use `open Syntax` and `open Paradigms` standard libraries. ### AI Agents * **Prompts:** Do not hardcode prompts in the agent logic. Import them from `ai_services/prompts.py`. * **Cost Control:** Ensure your agent logic respects the `MAX_RETRIES` defined in `client.py`. --- ## 4. Standards Compliance (Ninai & UD) ### Ninai Protocol * If you touch `app/adapters/ninai.py`, ensure you support the **Recursive Object Model**. * **Do not** revert to regex parsing. Use the `_walk_tree` recursion pattern. ### Universal Dependencies * If you add a new RGL function to `grammar_factory.py`, you **MUST** add its mapping to `app/core/exporters/ud_mapping.py`. * **Rule:** Every syntactic constructor must have a corresponding CoNLL-U tag map. --- ## 5. Quality Assurance (The Judge) We require **Gold Standard Validation** for all major PRs. 1. **Add Test Case:** Add a verified intent/text pair to `data/tests/gold_standard.json`. 2. **Run Judge:** `python -m pytest tests/integration/test_quality.py --use-judge`. 3. **Pass Criteria:** Your PR will be blocked if the Judge's similarity score drops below **0.8** for existing languages. --- ## 6. Commit Messages We follow the **Conventional Commits** specification: * `feat: add Zulu language support (Tier 3)` * `fix: resolve PGF overwriting bug in orchestrator` * `docs: update deployment guide for WSL 2` * `ai: optimize Architect Agent system prompt` * `test: add gold standard case for Hausa` --- ## 7. Hot-Reloading Note The `aw_worker` service watches the `semantik_architect.pgf` file. If you run a build, wait ~5 seconds for the logs to show: > `runtime_detected_file_change ... runtime_reloading_triggered` If this doesn't happen, check that your Docker volumes are correctly mounted to `/app`. ================================================================================================ FILE: docs/Technical-Reference/12-TOOLS_DASHBOARD_AND_API.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: ee22536e766d10c8b86fb2474fb373f02817bcf8e11b3615ea14f09b4063bcfd CONTENT_BYTES: 14854 ================================================================================================ # 🧰 Tools Dashboard & API **SemantiK Architect** This document defines the canonical architecture for the **Tools Dashboard** (`/semantik_architect/tools`) and the secure **Tools API** (`/api/v1/tools/*`). It covers: 1. The dashboard’s role in the system 2. The backend API contract 3. The security model 4. The workflow-oriented tool filtering model 5. The relationship between frontend registry metadata and backend execution policy 6. Validation and maintenance rules --- ## 1. Purpose The Tools Dashboard exists to provide a **safe, operator-friendly control surface** for running approved maintenance, build, QA, and diagnostics tasks without exposing arbitrary shell execution. The dashboard is not a generic terminal. It is a **curated orchestration UI** over a backend allowlist registry. Core goals: - expose only approved tool IDs - enforce argument policy - preserve repo-root confinement - provide consistent lifecycle and output logs - separate **normal workflows** from **power-user / debug tooling** - guide operators with **recommended workflows**, not just a flat inventory --- ## 2. Canonical Paths ### 2.1 Frontend route The Tools Dashboard is served at: ```text /semantik_architect/tools ```` ### 2.2 API base The Tools API is served under: ```text /semantik_architect/api/v1 ``` Canonical tools endpoints: ```text GET /api/v1/tools/registry POST /api/v1/tools/run ``` ### 2.3 Repo execution root All tool execution must be confined to: ```text FILESYSTEM_REPO_PATH ``` The backend must resolve tool targets relative to this repo root and reject paths outside it. --- ## 3. Design Principles ### 3.1 Workflow-first UI The dashboard must be organized around **operator intent**, not backend implementation categories. The primary selector is: ```text Workflow / Tool Set ``` This is the main navigation model for the page. ### 3.2 Power user is a visibility modifier **Power user** is **not** a workflow. It is a visibility switch that reveals hidden, debug, test, internal, heavy, or advanced tools within the currently selected workflow. ### 3.3 Curated normal mode Normal mode must show a curated subset of tools that support common workflows. Hidden tools, raw tests, scanners, and risky internals should remain out of sight unless Power user is enabled. ### 3.4 Recommended workflow guidance When a workflow filter is selected, the page must display a **Recommended workflow** card that explains the normal step order for that tool set. The dashboard should help users answer: * what should I run first? * what should I run next? * which tools are normal vs optional vs recovery-only? --- ## 4. Backend API Contract ## 4.1 `GET /api/v1/tools/registry` Returns the safe tool registry for the dashboard. Each entry must describe: * canonical `tool_id` * label * category * description * execution policy * allowed flags * positionals policy * hidden/internal/heavy/test metadata * workflow metadata used by the dashboard ### 4.1.1 Canonical metadata shape Recommended registry metadata shape: ```json { "tool_id": "language_health", "label": "Language Health", "category": "health", "description": "Compile + runtime health checks for one or more languages.", "hidden": false, "internal": false, "heavy": false, "is_test": false, "requires_ai_enabled": false, "allowed_flags": ["--mode", "--langs", "--json", "--verbose", "--fast"], "allow_positionals": false, "workflow_tags": ["recommended", "languageIntegration", "qaValidation"], "normal_path": true, "power_user_only": false, "recommended_order": 50 } ``` ### 4.1.2 Workflow metadata The backend registry should provide workflow-oriented metadata so the frontend does not need to hardcode all tool grouping logic. Recommended fields: * `workflow_tags: string[]` * `normal_path: boolean` * `power_user_only: boolean` * `recommended_order: number` These fields are presentation metadata only. They do **not** affect execution permissions. --- ## 4.2 `POST /api/v1/tools/run` Executes a registered tool using a safe request envelope. Request shape: ```json { "tool_id": "language_health", "args": ["--mode", "both", "--langs", "en", "fr", "--json", "--verbose"], "dry_run": false } ``` Response shape must include: * trace ID * lifecycle events * stdout * stderr * exit code * timing metadata * argument validation warnings/errors Example response envelope: ```json { "trace_id": "uuid", "started_at": "2026-03-07T13:07:04.674490Z", "duration_ms": 68743, "lifecycle": [ { "level": "INFO", "event": "request_received", "message": "Tool run request received" }, { "level": "INFO", "event": "tool_validated", "message": "Tool found in registry" }, { "level": "INFO", "event": "args_validated", "message": "All arguments accepted." }, { "level": "INFO", "event": "process_spawned", "message": "Executing command with timeout 1800s" }, { "level": "INFO", "event": "process_exited", "message": "Process exited with code 0" } ], "stdout": "...", "stderr": "...", "exit_code": 0 } ``` --- ## 5. Security Model The Tools API is **admin-only** and must never allow arbitrary command execution. Required protections: * strict allowlist registry * canonical tool IDs only * no aliases or remaps * repo-root confinement under `FILESYSTEM_REPO_PATH` * flag allowlisting * optional flag value-shape validation * per-tool timeout enforcement * output truncation * AI tool gating * authenticated access ### 5.1 Required rules * only tool IDs present in `TOOL_REGISTRY` may execute * only allowlisted flags may pass * disallowed positionals must be rejected * commands must execute with `cwd = FILESYSTEM_REPO_PATH` * tool targets resolving outside the repo root must be rejected * AI tools must return 403 unless explicitly enabled * output must be truncated to configured limits * the router must record lifecycle telemetry for every run --- ## 6. Dashboard UX Model ## 6.1 Main controls The dashboard should expose: * search input * **Workflow / Tool Set** dropdown * **Power user** checkbox * optional dry-run toggle * advanced filters (wired only, show tests, show internal, show heavy, show legacy) * health refresh action * visible/wired counts ### 6.1.1 Main workflow dropdown values Canonical dropdown values: * `recommended` * `languageIntegration` * `lexiconWork` * `buildMatrix` * `qaValidation` * `debugRecovery` * `aiAssist` * `all` Display labels: * Recommended * Language Integration * Lexicon Work * Build & Matrix * QA & Validation * Debug & Recovery * AI Assist * All Tools ### 6.1.2 Power user behavior When Power user is off: * hide debug-only tools * hide internal-only tools * hide raw test tools * hide heavy tools unless explicitly allowed * prefer curated normal-path tools When Power user is on: * reveal hidden/debug/internal/test/heavy tools according to the active workflow * expose advanced filters * preserve the current workflow selection --- ## 6.2 Recommended workflow card Selecting a workflow must show a card with: * workflow label * one-line goal * ordered steps * required tools * optional tools * warning if Power user is needed for parts of the flow This card is instructional only. It does not trigger execution automatically. --- ## 7. Canonical Workflow Bundles ## 7.1 Recommended **Goal:** shortest safe path for most operator tasks. Visible tools: * `build_index` * `compile_pgf` * `language_health` * `run_judge` Recommended workflow: 1. Build Index 2. Compile PGF 3. Language Health 4. Generate a sentence 5. Run Judge --- ## 7.2 Language Integration **Goal:** add, repair, and validate one language. Visible tools: * `build_index` * `lexicon_coverage` * `compile_pgf` * `language_health` * `run_judge` * `harvest_lexicon` * `gap_filler` * `bootstrap_tier1` Recommended workflow: 1. Add or change language files 2. Build Index 3. Lexicon Coverage 4. Harvest / Gap Fill if needed 5. Bootstrap Tier 1 if needed 6. Compile PGF 7. Language Health 8. Generate a sentence 9. Run Judge --- ## 7.3 Lexicon Work **Goal:** build or repair lexicon data. Visible tools: * `harvest_lexicon` * `gap_filler` * `lexicon_coverage` Power-user add-ons may include: * `seed_lexicon` * import/build helpers Recommended workflow: 1. Harvest or seed lexicon data 2. Fill gaps 3. Run Lexicon Coverage 4. Build Index 5. Run Language Health --- ## 7.4 Build & Matrix **Goal:** manage inventory/build state. Visible tools: * `build_index` * `compile_pgf` Power-user add-ons: * `rgl_scanner` * `lexicon_scanner` * `app_scanner` * `qa_scanner` * `bootstrap_tier1` Recommended workflow: 1. Build Index 2. Compile PGF 3. Language Health Power-user note: Use individual scanners only when the Everything Matrix looks wrong or stale. --- ## 7.5 QA & Validation **Goal:** verify correctness, runtime health, and performance. Visible tools: * `language_health` * `run_judge` * `profiler` Power-user add-ons: * raw smoke/API/GF/multilingual tests * regression generators Recommended workflow: 1. Language Health 2. Generate a sentence 3. Run Judge 4. Profiler --- ## 7.6 Debug & Recovery **Goal:** isolate broken or inconsistent states. Visible tools: * `diagnostic_audit` Power-user add-ons: * targeted scanners * raw pytest tools * low-level diagnostics Recommended workflow: 1. Diagnostic Audit 2. Run targeted scanner or raw test 3. Fix the issue 4. Build Index 5. Compile PGF 6. Language Health --- ## 7.7 AI Assist **Goal:** human-guided AI help for difficult gaps. Visible only when Power user is enabled: * `ai_refiner` * `seed_lexicon` Recommended workflow: 1. Confirm deterministic tools reveal a real gap 2. Use AI assistance 3. Build Index 4. Compile PGF 5. Language Health 6. Run Judge AI tools must never replace the deterministic normal path. --- ## 8. Tool Classification Rules The dashboard may still maintain internal categories such as: * build * maintenance * health * qa * data * ai * internal However, those categories are **secondary metadata**, not the primary user navigation model. The page must filter by **workflow first**, then optionally group or decorate by category. --- ## 9. Frontend Architecture Recommended frontend responsibilities: ### 9.1 `page.tsx` Owns: * tool data loading * filter state * workflow selection * visible tool list derivation * selected tool state * runner state * workflow card state ### 9.2 `useToolsPrefs.ts` Owns persisted user preferences, including: * `workflowFilter` * `powerUser` * `showLegacy` * `showTests` * `showInternal` * `wiredOnly` * `showHeavy` * layout preferences * console behavior preferences ### 9.3 `workflows.ts` Owns frontend workflow metadata: * workflow IDs and labels * fallback workflow card copy * fallback tool membership * fallback recommended order This file exists as a UI helper even when backend registry metadata is present. ### 9.4 `backendRegistry.ts` Owns curated frontend tool metadata and must remain in sync with the backend registry. ### 9.5 `buildToolItems.ts` Builds normalized dashboard tool items from backend registry + frontend presentation metadata. --- ## 10. Backend Architecture ## 10.1 Router `app/adapters/api/routers/tools.py` is the canonical tools router. It owns: * `/api/v1/tools/registry` * `/api/v1/tools/run` * request validation * argument policy enforcement * lifecycle envelope generation * timeout and truncation handling ## 10.2 Registry modules Execution policy lives in the backend registry modules, including: * build tools registry * maintenance tools registry * QA registry * AI / optional registry if present These registries define: * tool target * allowed flags * positionals policy * timeout * AI gating * hidden/internal metadata * workflow metadata for the frontend ## 10.3 Shared models The tools API models must explicitly support workflow metadata so frontend and backend stay aligned. Recommended model additions: * `workflow_tags: list[str]` * `normal_path: bool` * `power_user_only: bool` * `recommended_order: int | None` --- ## 11. Sync Rules The following must remain synchronized: * backend `TOOL_REGISTRY` * frontend `backendRegistry.ts` * workflow bundles shown in the dashboard * documentation in this file * documentation in `docs/17-TOOLS_AND_TESTS_INVENTORY.md` Any tool added to the backend registry should be reviewed for: 1. normal visibility 2. workflow membership 3. power-user status 4. parameter docs 5. documentation impact --- ## 12. Validation Checklist ## 12.1 Registry/API checks * `GET /api/v1/tools/registry` returns canonical tool IDs * hidden/internal/heavy/test metadata is present * workflow metadata is present * invalid tool IDs return 404 * invalid flags are rejected * AI tools return 403 when disabled ## 12.2 Dashboard checks * workflow dropdown changes visible tools * Power user reveals advanced tools without changing the active workflow * recommended workflow card updates correctly * `build_index` appears in normal workflows * advanced filters still work * counts update correctly * tool selection remains stable where possible during filter changes ## 12.3 Execution checks * `language_health` runs and shows lifecycle events * `diagnostic_audit` runs and shows argument validation * `lexicon_coverage` rejects unsupported flags * `compile_pgf` runs with allowed flags only * dry-run mode displays the resolved command without execution --- ## 13. Operational Guidance ### 13.1 Normal operator behavior Use workflow filters, not raw category browsing. Preferred order: * Recommended * Language Integration * QA & Validation ### 13.2 Power-user behavior Use Power user only when: * the normal workflow is insufficient * you need scanners or raw tests * you are debugging registry/matrix drift * you are working with AI-only tooling ### 13.3 Do not use the Tools Dashboard for * arbitrary shell execution * ad-hoc filesystem access outside repo policy * replacing deterministic build steps with AI calls * bypassing the official registry --- ## 14. Future Extensions Planned or supported future improvements: * registry-served workflow descriptions * per-workflow analytics * pinned favorite tools * workflow-specific presets * operator runbooks linked from workflow cards * richer tool dependency graph visualization --- ## 15. Summary The Tools Dashboard is a **workflow-oriented, allowlisted operator console** over the secure Tools API. The final model is: * **workflow dropdown** = what the user is trying to do * **Power user** = whether hidden/debug tools are visible * **registry metadata** = execution policy + presentation metadata * **recommended workflow card** = guidance for the selected tool set This keeps the page safe, scalable, and understandable even as the tool inventory grows. ================================================================================================ FILE: docs/Technical-Reference/12-WIKIMEDIA_ALIGNMENT.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: b9f16d02648618a644a27ea68749d3c3962e51cf297dee759258330a43dcf1f0 CONTENT_BYTES: 5144 ================================================================================================ # 🌐 Abstract Wikipedia Alignment & Standards **SemantiK Architect v2.0** This document defines how the **SemantiK Architect (SKA)** aligns with, diverges from, and integrates with the official technical standards of the **Abstract Wikipedia** project (Wikifunctions, Ninai, and Universal Dependencies). --- ## 1. The Core Divergence: Hybrid Linearization (GF + UD) The original Abstract Wikipedia architecture often relies heavily on **Universal Dependencies (UD)** for linguistic modeling (as seen in the `Udiron` project). SKA v1.0 was purely **Grammatical Framework (GF)** based. In **v2.0**, we have bridged this gap by adopting a **Hybrid Linearization Strategy**. ### 1.1 Tier 1: Generative Grammar (GF RGL) For high-resource languages (English, French, Hindi), we use the **GF Resource Grammar Library**. * **Why:** It handles complex morphology (case declension, gender agreement) perfectly using valid Abstract Syntax Trees. * **Alignment:** This provides the "Verifiable Correctness" required by the platform. ### 1.2 Tier 3: Weighted Topology (Udiron Integration) For under-resourced languages (Zulu, Hausa), we have integrated the **Weighted Topology** approach from the `Udiron` codebase. * **Why:** Writing full generative grammars for 300 languages is too slow. * **Mechanism:** We use `data/config/topology_weights.json` to define relative weights for dependency roles (e.g., `subj=-10`, `obj=-5`, `root=0` for SOV). * **Alignment:** This aligns SKA's "Factory" tier directly with the community's preferred method for rapid language expansion. ### 1.3 Evaluation: Universal Dependencies Export We acknowledge UD as the gold standard for *evaluation*. * **Feature:** SKA v2.0 supports `Accept: text/x-conllu`. * **Logic:** We implement **"Construction-Time Tagging."** Since we generate the sentence, we know exactly which word is the Subject. We map our internal RGL functions (`mkCl`) to UD tags (`nsubj`, `root`) using the **Frozen Ledger** mapping. * **Result:** SKA output can be validated against standard UD treebanks. --- ## 2. Ninai Protocol Integration **Ninai** is the abstract notation used by Abstract Wikipedia to represent meaning. * *Legacy Assumption:* LISP-like S-expressions. * *Code Reality:* Recursive JSON Object Trees (Constructors). ### 2.1 The Bridge (`NinaiAdapter`) SKA is designed to be a native **Renderer Implementation** for Ninai. * **Input:** We accept the recursive JSON structure natively. * **Mapping:** The `app/adapters/ninai.py` module recursively walks the Ninai tree and flattens it into SKA's internal `BioFrame` or `EventFrame`. ### 2.2 Constructor Mapping We map Ninai constructors to SKA logic: | Ninai Constructor | SKA Component | | --- | --- | | `ninai.constructors.Statement` | `BioFrame` (Root Intent) | | `ninai.constructors.List` | Recursive Flattening Logic | | `ninai.constructors.Entity` | `DiscourseEntity` (QID Lookup) | | `ninai.types.Bio` | `frame_type="bio"` | We view Ninai as the *wire format* (Z7) and SKA as the *execution engine* (Z1). --- ## 3. Z-Object Integration (Wikifunctions) In the Wikifunctions ecosystem, functions and types are assigned **Z-IDs**. SKA's architecture is "Z-Ready" by design. ### 3.1 Component Mapping | SKA Component | Wikifunctions Concept | Integration Strategy | | --- | --- | --- | | **Family Engine** (`RomanceEngine`) | **Z-Implementation** | Python code wrapped as a Z-Function. | | **Lexicon Entry** (`people.json`) | **Z-Object (Type)** | Mapped to `Z_Physicist` or Wikidata QIDs. | | **Matrix Config** (`por.json`) | **Z-Configuration** | Stored as a JSON Z-Object. | | **The Architect Agent** | **Z-Bot** | An automated contributor bot. | ### 3.2 Entity Grounding We utilize **Wikidata QIDs** (e.g., `Q42`) as the source of truth. The `NinaiAdapter` expects these QIDs in the `Entity` constructor arguments. --- ## 4. Discourse & Coherence Abstract Wikipedia aims to generate **Articles**, not just sentences. SKA v2.0 addresses this via the **Discourse Planner**. ### 4.1 Centering Theory We implement a simplified version of Centering Theory to handle **Pronominalization**. * **Standard:** If an entity is the "Backward-Looking Center" () of the previous utterance, it should be pronominalized. * **Implementation:** The `SessionContext` in Redis tracks the current focus. If the incoming Ninai frame references the same QID, SKA renders "She/He" instead of the name. --- ## 5. Addressing "Vibe-Coding" (Rigorous Engineering) To ensure this project is robust enough for the Wikimedia ecosystem, we enforce: 1. **Hexagonal Architecture:** Strict isolation of domain logic from the Ninai/UD adapters. 2. **Gold Standard QA:** We ingest the `Udiron` test suite (`tests.json`) to validate our outputs against community-verified strings. 3. **Two-Phase Compilation:** Solving the PGF linking bug deterministically. 4. **Self-Healing CI/CD:** The **Surgeon** and **Architect** agents automatically repair broken grammars, ensuring the build pipeline is resilient. We invite the community to review `docs/01-ENGINE_ARCHITECTURE.md` for a deep dive into these engineering standards. ================================================================================================ FILE: docs/Technical-Reference/15-SCHEMA_ALIGNMENT_PROTOCOL.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: e51434098626f157e829cdd1e3b15df7c8e90840b03a65bcca82641590b96bd2 CONTENT_BYTES: 4648 ================================================================================================ Here is the updated **`docs/15-SCHEMA_ALIGNMENT_PROTOCOL.md`**. This version incorporates the **v2.1 Overloading Strategy** (handling partial data via `mkBioFull`/`mkBioProf`) and the **WordNet/RGL integration** defined in your architecture specifications. --- # 📐 Schema Alignment Protocol & The "Triangle of Doom" **SemantiK Architect v2.1** ## 1. The Core Problem: Alignment Failure In the v2.1 architecture, a runtime error (`400 Bad Request` or `Function not found`) often occurs not because code is broken, but because the system's three definitions of "Truth" are out of sync. We call this the **Triangle of Doom**: 1. **The API Contract (Input):** What the user sends (JSON). Defined in `schemas/*.json` and `app/adapters/ninai.py`. 2. **The Abstract Grammar (Interface):** What the engine accepts (GF). Defined in `gf/semantik_architect.gf`. 3. **The Factory Logic (Generator):** What the builder produces (Python). Defined in `utils/grammar_factory.py`. ### The Symptom * **Error:** `Function 'mkBio' not found in grammar` or `unknown function` in compiler logs. * **Cause:** The API sent a `frame_type="bio"` (expecting `mkBio`), but the Abstract Grammar only defined `mkFact` or required different arguments. --- ## 2. The Solution: Manual Propagation Protocol Until a "Schema-to-Grammar" compiler is built, we explicitly adopt a **Manual Propagation Strategy**. To add or fix a Semantic Frame (e.g., `Event`, `Bio`, `Location`), you **MUST** update all three vertices of the triangle simultaneously. ### Step 1: Update the Interface (Abstract Grammar) **File:** `gf/semantik_architect.gf` Define the function signature. **v2.1 Mandate:** Use **Overloading** to handle missing data (e.g., when a user provides a Profession but no Nationality). ```haskell cat Statement ; Entity ; Profession ; Nationality ; fun -- The "Perfect" Case (All data available) mkBioFull : Entity -> Profession -> Nationality -> Statement ; -- The "Partial" Cases (Graceful Degradation) mkBioProf : Entity -> Profession -> Statement ; mkBioNat : Entity -> Nationality -> Statement ; -- Type Coercion (Bridge from WordNet types) lexProf : N -> Profession ; lexNat : A -> Nationality ; ``` ### Step 2: Update the Factory (Tier 3 Generation) **File:** `utils/grammar_factory.py` Teach the "Safe Mode" generator how to linearize these new functions for under-resourced languages (Tier 3). Since Tier 3 lacks the RGL's full power, we use **Weighted Topology** or simple concatenation stubs. ```python def generate_safe_mode_grammar(lang_code): # ... gf_code = f""" lin -- Tier 3 Stubs (Strings) mkBioFull name prof nat = name ++ "is a" ++ nat ++ prof; mkBioProf name prof = name ++ "is a" ++ prof; -- Coercion Stubs lexProf n = n; lexNat n = n; """ ``` ### Step 3: Update Tier 1 Concrete Grammars **File:** `gf/WikiEng.gf` (and other RGL languages) For High-Resource languages, you must link the **Abstract** functions to the **Concrete** RGL logic and the **WordNet** lexicon. ```haskell concrete WikiEng of SemantikArchitect = open SyntaxEng, ParadigmsEng, WordNetEng in { lincat Statement = S ; Profession = CN ; Nationality = AP ; lin -- Use RGL macros (mkS, mkCl, mkVP) for grammatically correct output mkBioFull s p n = mkS (mkCl s (mkVP n p)) ; -- "He is an American physicist" mkBioProf s p = mkS (mkCl s (mkVP (mkCN p))) ; -- "He is a physicist" -- Coercion lexProf n = mkCN n ; } ``` --- ## 3. Decision Record (ADR) ### Context The API layer (`NinaiAdapter`) is dynamic and handles optional JSON fields. The Grammar layer (`GF`) is static, strictly typed, and requires fixed arity (argument counts). ### Decision We choose **Explicit Semantic Mapping** with **Overloading** over **Generic Triples**. * **Option A (Rejected):** Use a generic `mkTriple : Subject -> Predicate -> Object -> Fact` for everything. * *Pros:* No need to update grammar for new frames. * *Cons:* Loses semantic nuance (e.g., "Born in" vs "Located in") required for accurate translations and UD tagging. * **Option B (Accepted):** Define specific functions `mkBioFull`, `mkBioProf`. * *Pros:* Allows language-specific handling (e.g., French uses "né en" for birth, "situé à" for location) and handles missing data gracefully. * *Cons:* Requires the 3-step manual update process described above. ### Future Roadmap To automate this, we will eventually implement an **Abstract Generator** script that reads `schemas/frames/*.json` and auto-generates `semantik_architect.gf` during the build pre-flight check. ================================================================================================ FILE: docs/Technical-Reference/16-DEV_TOOLS_AND_LAUNCHER.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 62e1dac74d4143f7e91f13f0225793961f3e7ce0645137e57d487be759209c88 CONTENT_BYTES: 5433 ================================================================================================ # 🛠️ Developer Tools & Unified Launch System **SemantiK Architect v2.5** ## 1. The "God Mode" Launcher (`Run-Architect.ps1`) The **Unified Orchestrator** is a PowerShell script that manages the hybrid environment (Windows Frontend + Linux/WSL Backend) automatically. **Location:** `Run-Architect.ps1` (Root) ### Why it exists 1. **Zombie killing:** It forcefully kills lingering `uvicorn` (Linux) or `node` (Windows) processes holding ports 8000 or 3000 to prevent address errors. 2. **Window Management:** It spawns **3 separate visible windows** for simultaneous monitoring of API, Worker, and Frontend logs. 3. **Consistency:** It delegates startup arguments to `manage.py`, ensuring the environment matches the system’s canonical configuration. ### Usage Right-click the file and select **"Run with PowerShell"**, or run from the terminal: ```powershell .\Run-Architect.ps1 ```` ### Process Flow * **Host (PowerShell):** Performs cleanup, verifies Docker/Redis, and spawns child windows. * **Terminal 1 (WSL):** API Backend via `python3 manage.py start-api`. * **Terminal 2 (WSL):** Background Worker via `python3 manage.py start-worker`. * **Terminal 3 (Windows Native):** Frontend via `npm run dev`. --- ## 2. The Developer Console (`/dev`) A dedicated dashboard for immediate system health verification and smoke testing. **URL:** `http://localhost:3000/semantik_architect/dev` ### Features * **System Heartbeat:** Real-time component status (ready/healthy) from the backend health endpoint. * **Broker:** Redis connection status. * **Storage:** Lexicon file accessibility. * **Engine:** PGF binary loading status. * **One-Click Smoke Test:** Sends a standard test payload to verify end-to-end generation without manual `curl`. * **Command Cheat Sheet:** Quick reference for common restart commands. --- ## 3. The System Tools Dashboard (`/tools`) A GUI wrapper for backend maintenance scripts, allowing operational tasks to be performed without a WSL terminal. **URL:** `http://localhost:3000/semantik_architect/tools` ### How it works (authoritative) 1. **Request:** The frontend sends a tool run request to `POST /api/v1/tools/run`: * `tool_id`: allowlisted identifier (e.g., `language_health`, `compile_pgf`) * `args`: optional argv-style list (backend validates/filters) * `dry_run`: if true, returns the resolved command without executing 2. **Validation:** The backend checks `tool_id` against a strict **Allowlist Registry** and validates flags/arg-shapes per-tool (prevents flag injection and arbitrary execution). 3. **Execution:** The backend spawns a subprocess **from the configured repo root** (`FILESYSTEM_REPO_PATH`) with environment injection (`PYTHONPATH`, `PYTHONUNBUFFERED`, `TOOL_TRACE_ID`). 4. **Result envelope:** The backend returns a stable response envelope including: * `trace_id`, `success`, `command` * `stdout`/`stderr` (plus back-compat `output`/`error`) * `exit_code`, `duration_ms` * `args_received` / `args_accepted` / `args_rejected` * `truncation` metadata (stdout/stderr) * `events` lifecycle telemetry (INFO/WARN/ERROR steps) ### Security & operational guarantees * **No arbitrary execution:** only allowlisted tool IDs can run. * **No aliases / no remaps:** tool IDs are canonical; legacy IDs are rejected (404 from the registry lookup). * **Repo confinement:** tool targets must resolve under `FILESYSTEM_REPO_PATH`. * **Timeouts:** per-tool timeout enforced; default via `ARCHITECT_TOOLS_DEFAULT_TIMEOUT_SEC`. * **Output truncation:** enforced via `ARCHITECT_TOOLS_MAX_OUTPUT_CHARS`. * **AI gating:** AI tools return 403 unless `ARCHITECT_ENABLE_AI_TOOLS=1`. * **Auth:** tools router is protected by API key (`verify_api_key`) and is treated as **admin-only**. ### Available Dashboard Tool Mappings (examples) | Tool ID | Action Taken | | ----------------- | ------------------------------------------------------------------- | | `language_health` | Language health/diagnostics utility (compile/API checks). | | `compile_pgf` | Triggers the build orchestrator to compile/link `semantik_architect.pgf`. | | `harvest_lexicon` | Runs the lexicon harvester (subcommands: `wordnet` or `wikidata`). | | `run_judge` | Executes golden-standard regression checks (AI Judge integration). | Note: the UI registry of tools (labels/categories/parameter docs) must stay in sync with backend `TOOL_REGISTRY` tool IDs and their allowed flags. --- ## 4. Backend Orchestration Details The developer interfaces above rely on the following infrastructure: * **Secure Tools Router:** `app/adapters/api/routers/tools.py` * Implements `/api/v1/tools/registry` and `/api/v1/tools/run` * Enforces allowlist registry + argument policy + truncation + timeouts + telemetry * **Tool Registry:** `TOOL_REGISTRY` maps safe tool IDs to physical scripts/tests and their execution policy: * allowed flags * whether positionals are permitted * flag value-shape rules (single-value vs multi-value) * AI gating (`requires_ai_enabled`) * **Repo Root Standardization:** tools always execute with `cwd = FILESYSTEM_REPO_PATH` and `PYTHONPATH` injected accordingly. * **Unified Commander:** `manage.py` remains the canonical orchestrator for lifecycle operations; external launchers should delegate to it rather than re-implement environment logic. ================================================================================================ FILE: docs/Technical-Reference/17-TOOLS_AND_TESTS_INVENTORY.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 2d2fc021339e9b90c3dbf3e7c213c347bef9477f388440975c5503a5bfcbd958 CONTENT_BYTES: 36448 ================================================================================================ # 📚 The Complete Tools & Tests Inventory (Final) **SemantiK Architect** Status: normative Owner: Tools / QA / Runtime / Frontend Scope: complete inventory and usage model for tools, test surfaces, QA flows, and AI-assisted utilities in the final planner-first multilingual runtime This document is the **Single Source of Truth** for: 1. **GUI Tools** (Web Dashboard) 2. **Workflow Filters** (Tools Page) 3. **CLI Orchestration** (Backend Management) 4. **Build & Matrix Operations** 5. **Diagnostics & Recovery** 6. **Data Operations** (Lexicon & Imports) 7. **Quality Assurance** (Testing & Validation) 8. **AI Services** (Agents and AI-gated tools) 9. **Pytest Surfaces** (Regression and acceptance tests) It exists to prevent the following failure modes: * treating debug-only tools as part of the normal path, * confusing matrix/scanner tools with user-facing workflows, * treating compile success as language readiness, * treating non-empty output as acceptance success, * and letting EN/FR cutover validation drift away from the final planner-first runtime model. --- ## 0. Architectural anchor All tooling and testing described here must align with the final runtime model: ```text canonical input -> normalized frame/domain form -> planner -> lexical resolution -> realizer -> SurfaceResult -> public response mapping -> HTTP JSON response ``` The nominal runtime is **planner-first**. Normal tool and test flows must validate: * planner-first runtime behavior, * explicit runtime metadata, * correct concrete language realization, * stable public response contract, * and language readiness beyond mere compilation or routing. For EN/FR bio/person generation, the accepted vertical slice is: * request normalization, * planner-first generation, * language-specific realization, * coherent public envelope, * acceptance checks, * and surface-language correctness. --- ## 1. Tools Dashboard UX model The Tools Dashboard is organized around **user intent**, not backend internals. ### Main controls * **Workflow / Tool Set** dropdown * **Power user (debug)** checkbox * **Advanced filters** shown only when Power user is enabled ### Workflow filters | Workflow Filter | Purpose | Normal Tool Set | | ------------------------ | ------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------ | | **Recommended** | Short deterministic path for most work | `build_index`, `compile_pgf`, `language_health`, `run_judge` | | **EN/FR Cutover** | Final runtime/API/QA validation for EN/FR bio/person | `build_index`, `compile_pgf`, `language_health`, `eval_bios`, `run_judge` | | **Language Integration** | Add or repair one language | `build_index`, `lexicon_coverage`, `compile_pgf`, `language_health`, `run_judge`, `harvest_lexicon`, `gap_filler`, `bootstrap_tier1` | | **Lexicon Work** | Data and vocabulary work | `harvest_lexicon`, `gap_filler`, `lexicon_coverage` | | **Build & Matrix** | Build-state / inventory work | `build_index`, `compile_pgf` | | **QA & Validation** | Runtime, contract, regression, acceptance, performance | `language_health`, `eval_bios`, `run_judge`, `profiler` | | **Debug & Recovery** | Broken or inconsistent system state | `diagnostic_audit` | | **AI Assist** | AI-gated repair or bootstrap flows | `ai_refiner`, `seed_lexicon_ai` | | **All** | Full visible inventory | All visible tools | ### Power user behavior **Power user** is a **visibility modifier**, not a workflow. When enabled, it may reveal: * hidden tools * scanner-level tools * test-oriented tools * internal tools * heavy tools * legacy or transitional tools ### Recommended workflow cards When a workflow filter is selected, the UI should display a short **Recommended Workflow** card. Examples: * **Recommended** `Build Index → Compile PGF → Language Health → Generate sentence → Run Judge` * **EN/FR Cutover** `Change code/docs → Build Index → Compile PGF → Language Health → Generate EN/FR planner-first examples → Run eval_bios → Run Judge` * **Language Integration** `Add/change files → Build Index → Lexicon Coverage → Harvest / Gap Fill if needed → Bootstrap Tier 1 if needed → Compile PGF → Language Health → Generate sentence → Run Judge` * **Lexicon Work** `Harvest / Seed → Gap Fill → Lexicon Coverage → Build Index → Language Health` * **Build & Matrix** `Build Index → Compile PGF → Language Health` * **QA & Validation** `Language Health → Generate sentence → Run eval_bios or Judge → Profiler` * **Debug & Recovery** `Diagnostic Audit → targeted scanner or pytest surface → fix → Build Index → Compile PGF → Language Health` * **AI Assist** `Use only after deterministic tools show a real gap → AI assist → Build Index → Compile PGF → Language Health → Run Judge` --- ## 2. Core orchestration Primary entry points for managing the overall system lifecycle. | Command / Script | Location | Purpose | Key Arguments | | ----------------------- | -------- | -------------------------------------------------------------------------- | ----------------------------------- | | **`manage.py`** | `Root` | Unified CLI for starting, building, and cleaning the system. | `start`, `build`, `doctor`, `clean` | | **`Run-Architect.ps1`** | `Root` | Windows launcher that handles process cleanup and starts the hybrid stack. | none | | **`Makefile`** | `Root` | Legacy build shortcuts and convenience wrappers. | `all`, `clean` | | **`StartWSL.bat`** | `Root` | Quick shell launcher into WSL with venv activated. | none | ### Core orchestration rule Normal runtime validation is not complete until it includes: 1. build or rebuild as needed, 2. runtime health checks, 3. at least one real generation request, 4. and the relevant QA or acceptance surface. --- ## 3. The build system Scripts that turn source grammars into runtime artifacts. | Tool | Location | Purpose | Key Arguments | | ----------------------- | ------------------------- | ------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | **Orchestrator** | `builder/orchestrator/` | Canonical build pipeline. Compiles intermediates and links the final PGF. | `--strategy`, `--langs`, `--clean`, `--verbose`, `--max-workers`, `--no-preflight`, `--regen-safe` | | **Orchestrator (Shim)** | `builder/orchestrator.py` | Backwards-compatible wrapper for legacy callers. | delegates to package entrypoint | | **Compiler** | `builder/compiler.py` | Low-level wrapper around `gf`. Manages includes and isolation. | internal | | **Strategist** | `builder/strategist.py` | Chooses build strategy and writes build plan. | internal | | **Forge** | `builder/forge.py` | Writes or materializes concrete grammar files according to build plan. | internal | | **Healer** | `builder/healer.py` | Reads build failures and dispatches AI repair for broken grammars. | internal | ### Build rule A successful build is not equivalent to language readiness. Build success proves only part of the stack: * grammar compiles, * artifacts link, * runtime artifacts exist. It does **not** prove: * planner-first correctness, * surface-language correctness, * public contract correctness, * or acceptance readiness. --- ## 4. The Everything Matrix System intelligence layer that scans repository state and language readiness signals. > **Important:** `build_index.py` is the **normal** entrypoint. Scanner scripts are Power user / debug tools unless a workflow explicitly requires them. | Tool | Location | Purpose | Key Arguments | | ------------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | | **Matrix Builder** | `tools/everything_matrix/build_index.py` | Scans RGL, Lexicon, App, and QA layers to build `everything_matrix.json`. Computes maturity signals and build strategies. | `--out`, `--langs …`, `--force`, `--regen-rgl`, `--regen-lex`, `--regen-app`, `--regen-qa`, `--verbose` | | **RGL Scanner** | `tools/everything_matrix/rgl_scanner.py` | Audits `gf-rgl/src` presence and consistency. | scanner-specific | | **Lexicon Scanner** | `tools/everything_matrix/lexicon_scanner.py` | Scores lexicon maturity by scanning coverage. | scanner-specific | | **App Scanner** | `tools/everything_matrix/app_scanner.py` | Scans backend/frontend surfaces for language support signals. | scanner-specific | | **QA Scanner** | `tools/everything_matrix/qa_scanner.py` | Parses QA artifacts and logs to update quality scoring. | scanner-specific | ### Matrix rule in normal workflows For onboarding and cutover work: 1. add or change files, 2. refresh the Everything Matrix, 3. validate build and runtime, 4. validate the public contract, 5. validate acceptance. ### Readiness rule Matrix presence or routing signals do not imply language readiness. A language is not acceptance-ready because it: * appears in inventory, * compiles, * loads, * routes, * or emits non-empty output. --- ## 5. Maintenance & diagnostics Tools used to keep the repository sane and the system healthy. > **GUI note:** The Tools Dashboard runs through a strict backend allowlist. The “Key Arguments” below reflect allowlisted argv flags for GUI execution. > **Security note:** Do **not** pass secrets via argv. Tool args may appear in logs, telemetry, UI output, or debug bundles. Use environment-based secret injection. | Tool | Location | Purpose | Key Arguments | Typical Workflow | | -------------------- | --------------------------- | -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | | **Language Health** | `tools/language_health.py` | Deep scan utility for language pipeline health. | `--mode`, `--fast`, `--parallel`, `--api-url`, `--timeout`, `--limit`, `--langs …`, `--no-disable-script`, `--verbose`, `--json` | Recommended, EN/FR Cutover, Language Integration, QA & Validation | | **Diagnostic Audit** | `tools/diagnostic_audit.py` | Forensics audit for stale artifacts and inconsistent outputs. | `--verbose`, `--json` | Debug & Recovery | | **Root Cleanup** | `tools/cleanup_root.py` | Moves loose artifacts into expected folders and cleans known junk outputs. | `--dry-run`, `--verbose`, `--json` | Debug & Recovery | | **Bootstrap Tier 1** | `tools/bootstrap_tier1.py` | Scaffolds Tier 1 wrappers or bridge files for selected languages. | `--langs …`, `--force`, `--dry-run`, `--verbose` | Language Integration | ### Health rule `language_health` is necessary but not sufficient for final acceptance. It is a health tool, not the entire proof layer. --- ## 6. Data operations Lexicon mining, harvesting, syncing, and vocabulary maintenance. | Tool | Location | Purpose | Key Arguments | Typical Workflow | | ---------------------------------------- | -------------------------------------- | ------------------------------------------------------------------------ | -------------------------------------------------------------------------------- | ---------------- | | **Universal Lexicon Harvester** | `tools/harvest_lexicon.py` | Two-mode harvester for lexicon data. | `wordnet ...`, `wikidata ...` | | | **Wikidata Importer (Legacy/Reference)** | `scripts/lexicon/wikidata_importer.py` | Legacy/reference importer logic. Not the authoritative v2 runtime path. | varies | | | **RGL Syncer** | `scripts/lexicon/sync_rgl.py` | Extracts lexical functions from compiled PGF into language shards. | `--pgf`, `--out-dir`, `--langs`, `--max-funs`, `--dry-run`, `--validate` | | | **Gap Filler** | `tools/lexicon/gap_filler.py` | Compares target lexicon vs pivot language to find missing concepts. | `--target`, `--pivot`, `--data-dir`, `--json-out`, `--verbose` | | | **Link Libraries** | `link_libraries.py` | Ensures `Wiki*.gf` opens required modules for runtime lexicon injection. | none | | | **Schema/Index Utilities** | `utils/...` | Lexicon index/schema and stats maintenance. | `refresh_lexicon_index.py`, `migrate_lexicon_schema.py`, `dump_lexicon_stats.py` | | | **Seed Lexicon (AI)** | `utils/seed_lexicon_ai.py` | Generates seed lexicon for selected languages. | AI-gated | | ### Lexicon rule Lexicon sufficiency is part of readiness, but it is still not enough by itself. A language with good lexicon coverage may still fail: * planner-first generation, * public contract parity, * concrete language realization, * or acceptance correctness. --- ## 7. Quality assurance tools QA tools that validate runtime output, contract shape, lexicon integrity, acceptance gates, and regression behavior. | Tool | Location | Purpose | Key Arguments | Typical Workflow | | ------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | ---------------------------------- | | **Universal Test Runner** | `tools/qa/universal_test_runner.py` | Runs CSV-based suites and emits a report. | `--suite`, `--in`, `--out`, `--langs …`, `--limit`, `--verbose`, `--fail-fast`, `--strict` | QA & Validation | | **Bio Evaluator** | `tools/qa/eval_bios.py` | EN/FR and multilingual bio/person evaluator. Validates response contract, runtime path, fallback behavior, language plausibility, and acceptance gates. | `--langs …`, `--limit`, `--out`, `--verbose` | EN/FR Cutover, QA & Validation | | **Lexicon Coverage Report** | `tools/qa/lexicon_coverage_report.py` | Coverage report for intended vs implemented lexicon and errors. | `--lang`, `--include-files`, `--verbose`, `--fail-on-errors` | Language Integration, Lexicon Work | | **Ambiguity Detector** | `tools/qa/ambiguity_detector.py` | Checks curated ambiguous sentences for multiple parse trees. | `--lang`, `--sentence`, `--topic`, `--json-out`, `--verbose` | QA & Validation | | **Batch Test Generator** | `tools/qa/batch_test_generator.py` | Generates large regression datasets for QA. | `--langs …`, `--out`, `--limit`, `--seed`, `--verbose` | QA & Validation | | **Test Suite Generator** | `tools/qa/test_suite_generator.py` | Generates empty CSV templates for manual fill-in. | `--langs …`, `--out`, `--verbose` | QA & Validation | | **Lexicon Regression Test Generator** | `tools/qa/generate_lexicon_regression_tests.py` | Builds lexicon regression tests for CI. | `--langs …`, `--out`, `--limit`, `--verbose`, `--lexicon-dir` | QA & Validation | | **Profiler** | `tools/health/profiler.py` | Benchmarks grammar/runtime performance. | `--lang`, `--iterations`, `--update-baseline`, `--threshold`, `--verbose` | QA & Validation | | **AST Visualizer** | `tools/debug/visualize_ast.py` | Generates JSON AST from sentence, intent, or explicit AST. | `--lang`, `--sentence`, `--ast`, `--pgf` | Debug & Recovery | ### EN/FR acceptance validation chain For the final EN/FR bio/person vertical slice: 1. `build_index` 2. `compile_pgf` 3. `language_health` 4. generate real EN and FR planner-first requests 5. validate the public response envelope 6. run `eval_bios` 7. run targeted pytest surfaces 8. only then treat EN/FR as accepted ### Evaluator rule `eval_bios` must reject false positives such as: * routed language looks correct but surface language is wrong, * planner-first claimed but required metadata is missing, * fallback used when nominal planner-first success is required. --- ## 8. AI services Autonomous agents and AI-gated tools. | Agent / Tool | File | Role | Triggered By | | --------------------- | ------------------------------ | ------------------------------------------------------------------------- | ---------------------------- | | **The Architect** | `ai_services/architect.py` | Generates missing grammars based on topology constraints. | Build/CLI workflow | | **The Surgeon** | `ai_services/surgeon.py` | Repairs broken `.gf` files using compiler logs. | `builder/healer.py` | | **The Lexicographer** | `ai_services/lexicographer.py` | Bootstraps core vocabulary for empty languages. | CLI / missing-data workflows | | **The Judge** | `ai_services/judge.py` | Grades generated text against gold standards and regression expectations. | quality workflows | | **AI Refiner** | `tools/ai_refiner.py` | Upgrades weak grammars toward RGL compliance. | AI-gated tools runner | | **Seed Lexicon (AI)** | `utils/seed_lexicon_ai.py` | Generates seed lexicon for selected languages. | AI-gated tools runner | ### AI gating Backend enforces `ARCHITECT_ENABLE_AI_TOOLS=1` for AI-gated tools. ### AI usage rule AI tools do not replace deterministic proof. They belong to **AI Assist** and should not appear in the normal deterministic path unless explicitly requested or Power user mode is enabled. --- ## 9. Pytest surfaces Automated regression and acceptance harness. Run with `pytest `. | Category | File | Description | | ------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | **Core** | `tests/core/test_use_cases.py` | Tests use cases such as `GenerateText`, including planner-first behavior, fallback behavior, and runtime metadata expectations. | | **Core** | `tests/core/test_domain_models.py` | Tests runtime model behavior including `SurfaceResult` expectations and debug/top-level consistency rules. | | **Integration** | `tests/integration/test_generate_via_planner_en.py` | EN planner-first generation integration checks. | | **Integration** | `tests/integration/test_generate_via_planner_fr.py` | FR planner-first generation integration checks. | | **Integration** | `tests/integration/test_quality.py` | Judge-based regression checks and quality evaluation. | | **Integration** | `tests/integration/test_worker_flow.py` | Verifies worker compilation/job flow. | | **Integration** | `tests/integration/test_ninai.py` | Tests Ninai adapter parsing logic. | | **Smoke** | `tests/test_api_smoke.py` | Checks `/health` and generation-related endpoints. | | **Smoke** | `tests/test_gf_dynamic.py` | Validates dynamic loading and linearization of GF grammars. | | **Smoke** | `tests/test_lexicon_smoke.py` | Validates lexicon JSON schema and syntax. | | **Multilingual** | `tests/test_multilingual_generation.py` | Cross-language generation regression surface. | | **Lexicon** | `tests/test_lexicon_loader.py` | Tests lazy-loading of lexicon shards. | | **Lexicon** | `tests/test_lexicon_index.py` | Tests in-memory indexing and lookups. | | **Lexicon** | `tests/test_lexicon_wikidata_bridge.py` | Tests Wikidata QID extraction and bridge logic. | | **Frames** | `tests/test_frames_*.py` | Unit tests for semantic frame dataclasses. | | **API** | `tests/http_api/test_generate.py` | Tests `POST /generate` request/response behavior. | | **API** | `tests/http_api/test_generations.py` | Tests generation API variants and response contract behavior. | | **API** | `tests/http_api/test_ai.py` | Tests AI suggestion endpoints. | | **Planning** | `tests/unit/planning/test_construction_plan.py` | Tests construction plan behavior and invariants. | | **Planning** | `tests/unit/planning/test_frame_to_plan.py` | Tests frame-to-plan conversion. | | **Planning** | `tests/unit/planning/test_frame_to_slots.py` | Tests slot extraction and mapping. | | **Lexicon Runtime** | `tests/unit/lexicon/test_lexical_resolution.py` | Tests lexical resolution behavior. | | **Use Cases** | `tests/unit/use_cases/test_plan_text.py` | Tests planning use case behavior. | | **Use Cases** | `tests/unit/use_cases/test_realize_text.py` | Tests realization use case behavior. | | **Renderers** | `tests/unit/renderers/test_family_construction_adapter.py` | Tests family renderer adapter behavior. | | **Renderers** | `tests/unit/renderers/test_gf_construction_adapter.py` | Tests GF renderer adapter behavior. | ### Pytest rule For final EN/FR cutover proof, pytest must cover: * nominal planner-first success, * explicit fallback behavior, * fallback-disabled failure behavior, * runtime metadata presence, * public contract expectations, * EN/FR realization correctness, * and regression against routed-but-wrong-language false positives. ### Pytest tools in the dashboard Pytest-backed tools belong mainly to: * **QA & Validation** * **Debug & Recovery** * Power user mode --- ## 10. Tools runner (backend API) The GUI runs tools through a strict backend allowlist registry. There is no arbitrary execution. | Endpoint | Purpose | | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `GET /api/v1/tools/registry` | Returns tool metadata, availability, UI metadata, and workflow metadata for the dashboard. | | `POST /api/v1/tools/run` | Runs a tool by `tool_id` plus argv-style args and optional dry-run mode. Returns a stable response envelope containing trace, command, stdout/stderr, truncation info, accepted/rejected args, lifecycle events, and exit code. | ### Request shape * `tool_id`: string * `args`: string[] * `dry_run`: boolean, optional ### Dry-run note Prefer using `dry_run=true` at the API layer instead of relying on per-tool argv conventions. ### Secret handling Do **not** pass API keys, tokens, or passwords in `args`. Use environment variables or secret injection instead. ### Execution constraints * repository root fixed by `FILESYSTEM_REPO_PATH` * output truncation by `ARCHITECT_TOOLS_MAX_OUTPUT_CHARS` * default timeout by `ARCHITECT_TOOLS_DEFAULT_TIMEOUT_SEC` * AI gating by `ARCHITECT_ENABLE_AI_TOOLS` --- ## 11. Registry metadata model The tools registry carries both execution metadata and UI/workflow metadata. ### Tool-level registry metadata Each tool may expose: * `tool_id` * `label` * `description` * `timeout_sec` * `allow_args` * `requires_ai_enabled` * `available` * `category` * `hidden` * `legacy` * `internal` * `heavy` * `is_test` * `allowed_flags` * `allow_positionals` * `flags_with_value` * `flags_with_multi_value` * `workflow_tags` * `workflow_order` ### Workflow registry metadata The registry may also expose a `workflows` array, where each workflow includes: * `workflow_id` * `label` * `summary` * `steps` * `tool_ids` * `power_user_addons` ### Why this exists This lets the frontend: * render the workflow dropdown, * render the recommended workflow card, * filter tools by user intent, * keep workflow taxonomy synchronized with backend truth. --- ## 12. Normal workflow reference ### Recommended 1. `build_index` 2. `compile_pgf` 3. `language_health` 4. generate a sentence 5. `run_judge` ### EN/FR Cutover 1. change grammar, runtime, contract, or docs 2. `build_index` 3. `compile_pgf` 4. `language_health` 5. generate real EN and FR planner-first requests 6. inspect public response shape 7. `eval_bios` 8. run targeted pytest surfaces 9. only then declare the slice accepted ### Language Integration 1. add or change language files 2. `build_index` 3. `lexicon_coverage` 4. `harvest_lexicon` or `gap_filler` if needed 5. `bootstrap_tier1` if needed 6. `compile_pgf` 7. `language_health` 8. generate a sentence 9. `run_judge` ### Lexicon Work 1. `harvest_lexicon` or `seed_lexicon_ai` 2. `gap_filler` 3. `lexicon_coverage` 4. `build_index` 5. `language_health` ### Build & Matrix 1. `build_index` 2. `compile_pgf` 3. `language_health` ### QA & Validation 1. `language_health` 2. generate a sentence 3. `eval_bios` or `run_judge` 4. run targeted pytest surface 5. `profiler` if needed ### Debug & Recovery 1. `diagnostic_audit` 2. targeted scanner or pytest surface 3. fix 4. `build_index` 5. `compile_pgf` 6. `language_health` ### AI Assist 1. confirm deterministic workflow failed or is incomplete 2. run AI assist tool 3. `build_index` 4. `compile_pgf` 5. `language_health` 6. `run_judge` --- ## 13. Acceptance-oriented rules ### 13.1 Build is not acceptance A language is not accepted because it: * compiles, * loads, * routes, * or emits non-empty text. ### 13.2 Planner-first is the nominal path The expected primary runtime is planner-first. Legacy success does not count as nominal planner-first acceptance. ### 13.3 EN/FR are the first full vertical slice EN and FR bio/person generation are the first acceptance-ready slice that must prove: * planner-first runtime, * correct concrete language realization, * coherent public contract, * evaluator success, * and regression protection. ### 13.4 Evaluators and tests must reject false positives The system must fail EN/FR acceptance when: * FR resolves correctly but surfaces English, * planner-first is claimed but required metadata is absent, * fallback is used where nominal success is required, * or top-level public fields and debug info disagree. --- ## 14. Summary rules * **Power user** is a visibility switch, not a workflow. * **Workflow dropdown** is the primary navigation model for the Tools page. * **`build_index`** is a normal visible workflow tool, not a debug-only tool. * **Scanners** are debug-level tools unless explicitly needed. * **AI tools** belong to **AI Assist**, not the normal deterministic path. * **Build success is not language readiness.** * **Routing success is not language correctness.** * **A language is not truly integrated until it generates correctly on the required runtime path and passes the relevant acceptance gates.** ================================================================================================ FILE: docs/Technical-Reference/A-SIMPLE-language_integration_workflow_reference.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: f0868cd30c1b8d87447c6297e4805cfaefd3b9e47908eb8525a92dc0ff67832e CONTENT_BYTES: 6962 ================================================================================================ **Language Integration Workflow** Quick Reference Purpose: the normal deterministic path to add or repair one language without drifting away from the runtime contract. | Core rule: change source files first, then refresh the Everything Matrix. The matrix is the system snapshot and status view, not the source of edits. | | :---------------------------------------------------------------------------------------------------------------------------------------------------- | # **Normal workflow** ## **1. Add or change language files** Put grammar, config, and lexicon files in place first. Do not treat generated artifacts or matrix output as the source of truth. ## **2. Refresh the Everything Matrix** Rebuild the matrix so the repo re-discovers the language, rebuild strategy, lexicon status, app status, and QA status. This step is required before trusting later compile or health results. ## **3. Validate the lexicon** Run lexicon validation before compile. The goal is to catch thin, empty, malformed, or structurally unusable lexical data early, before build and runtime checks. ## **4. Fill gaps only if needed** If coverage is too thin, use deterministic lexicon tools to import or fill gaps. These are repair tools, not part of the shortest normal path when the lexicon is already good enough. ## **5. Bootstrap Tier 1 scaffolding only if needed** Use Tier 1 bootstrap only when the language is missing required scaffolding. Do not run it as a routine step for languages that already have working source files. ## **6. Compile the PGF** Compile after the matrix refresh and lexicon validation. This is the point where the language must enter the grammar binary cleanly. ## **7. Validate compile and runtime** Run language health in both modes so compile and runtime are checked together. This is where you confirm that the language is not only buildable but callable through the runtime. ## **8. Generate one real sentence** Use the dev smoke path or call `/api/v1/generate/` and verify a real surface result comes back. Do not stop at “some text appeared”. Confirm at minimum: * non-empty `text` * correct `lang_code` * visible `renderer_backend` * explicit `fallback_used` * usable `debug_info` for runtime diagnosis ## **9. Confirm runtime path explicitly** When you inspect `debug_info`, verify which runtime path actually ran. The target architecture is planner-first; if the language is still served through the legacy direct path, that must stay explicit and treated as compatibility state, not hidden success. ## **10. Stabilize** Run judge or regression only after the language builds and generates. Then refresh the matrix again so final build and QA status are recorded in the system snapshot. # **Exact tool calls** | Step | Command | | | :--------------------------------- | :------------------------------------------------------------------------------------------------------------------------------ | - | | Refresh matrix | `build_index --langs --regen-rgl --regen-lex --regen-app --regen-qa --verbose` | | | Validate lexicon | `lexicon_coverage --lang --include-files` | | | Fill lexical gaps (only if needed) | `harvest_lexicon ...` or `gap_filler --langs --pivot en --verbose` | | | Bootstrap Tier 1 (only if needed) | `bootstrap_tier1 --langs --verbose` | | | Compile PGF | `compile_pgf --langs --verbose` | | | Validate health | `language_health --mode both --langs --json --verbose` | | | Generate sentence | Use Dev smoke test or call `/api/v1/generate/` | | | Stabilize | `run_judge --langs --verbose` then `build_index --langs --regen-rgl --regen-lex --regen-app --regen-qa --verbose` | | # **Required vs optional** | Required in the normal path | Use only when needed | | | :---------------------------------- | :------------------------------------------------ | - | | `build_index` | `run_judge` | | | `lexicon_coverage` | `harvest_lexicon` | | | `compile_pgf` | `gap_filler` | | | `language_health` | `bootstrap_tier1` | | | one real generation test | `diagnostic_audit` | | | runtime-path check via `debug_info` | low-level scanners and pytest-only recovery tools | | # **Dependencies to remember** * Use ISO-2 language codes in tool calls, for example `en`, `fr`, `pt`. * Refresh the matrix after source changes because later compile and status decisions depend on that snapshot. * Validate the lexicon before compile so empty or broken lexical shards are caught early. * Compile before final health validation so runtime checks are not reading stale binaries. * A language is not truly integrated until it generates a sentence through the runtime. * For migrated constructions, the authoritative target is planner-first orchestration with shared runtime contracts; legacy direct generation may still exist during migration, but it must remain explicit in diagnostics. * A successful generation check should validate the runtime result shape, not only the presence of text. The stable top-level diagnostics are `construction_id`, `renderer_backend`, and `fallback_used`, with additional detail in `debug_info`. # **Definition of done for one language** A language can be considered integrated for the normal path when all of the following are true: * source files are in place * the matrix sees the language correctly * lexicon validation passes * PGF compile passes * runtime health passes * one real generation succeeds for `/api/v1/generate/` * the result exposes valid runtime diagnostics * any fallback or legacy-path usage is explicit, not hidden For stronger acceptance on core constructions, add regression and judge coverage after the normal path is green. English and French planner-first integration coverage is part of the broader migration target, not just a nice-to-have. ================================================================================================ FILE: docs/Technical-Reference/Abstract_Wiki_Architect_Build_and_Launch_System.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 838853ba8de6bc9ff9a34493339e3a341671965e6b0072c27d3ed4aa43ff6240 CONTENT_BYTES: 4746 ================================================================================================ # 🚀 SemantiK Architect: Unified Build & Launch System (v2.0) ## 1. High-Level Architecture The system follows a strict **"Check, Build, Serve"** pipeline. It does not simply start the API; it first performs a deep census of your data (The Matrix), compiles the grammar binary (The PGF) using a robust two-phase process, and only then launches the runtime services. The v2.0 architecture replaces fragile OS-scripts with a **Unified Python Task Runner** to ensure speed, atomicity, and portability. --- ## 2. Level 1: The Commander (`manage.py`) 🚀 **Script:** `manage.py` (Replaces `launch.ps1`) **Role:** The Unified CLI & Environment Manager **Location:** Root This is the single entry point for all developer operations. It centralizes logic previously scattered across `.bat`, `.sh`, and `.ps1` files. ### Core Commands * **`python manage.py start`**: The "Daily Driver". Checks Docker/Redis, runs an incremental build, and launches the API/Worker. * **`python manage.py build --clean --parallel 8`**: Forces a clean rebuild using parallel processing. * **`python manage.py doctor`**: Runs system diagnostics (identifies Zombie files, checks pathing). * **`python manage.py generate --missing`**: Asynchronously calls AI/Factory to create missing grammars (Decoupled from build). --- ## 3. Level 2: The Optimized Builders 🏗️ These scripts run sequentially under the `build` command. They are refactored for **Speed** (Caching/Parallelism) and **Safety** (Atomic Writes). ### A. The Census Taker (Indexer) * **Script:** `tools/everything_matrix/build_index.py` * **Optimization:** **Content Hashing (Caching)**. * **Logic:** 1. Checks the MD5 checksum of `gf-rgl/` and `data/lexicon/`. 2. **Match:** Loads cached `everything_matrix.json` (0s latency). 3. **Mismatch:** Rescans the file system and updates the index. ### B. The Compiler (Orchestrator) * **Script:** `builder/orchestrator.py` * **Optimization:** **Parallel Verification & Single-Shot Linking**. * **Logic:** 1. **Weighted Topology (Pre-Flight):** Calls `grammar_factory.py` to generate "Safe Mode" grammars for Tier 3 languages (e.g., Zulu, Hausa) using `topology_weights.json`. 2. **Parallel Verification:** Spawns a process pool (e.g., 8 workers) to compile `.gfo` files for all languages simultaneously. 3. **Atomic Writes:** Compiles to a `_temp` directory first. Only valid builds are moved to the final folder, permanently eliminating "Zombie" files. 4. **Single-Shot Link:** Collects *all* valid languages and executes **one single** `gf -make` command to produce the final `semantik_architect.pgf` binary. --- ## 4. Level 3: The Specialists (Hidden Dependencies) 🧠 These are libraries imported by Level 2 to perform complex analysis, generation, or repair. ### A. The Weighted Topology Factory * **Script:** `utils/grammar_factory.py` * **Role:** Deterministic Generation (Tier 3). * **Logic:** Uses `data/config/topology_weights.json` to generate grammatically correct word order (SVO, SOV, VSO) for under-resourced languages without needing AI. ### B. The AI Agent (Architect & Surgeon) * **Script:** `ai_services/architect.py` * **Role:** Probabilistic Generation (Tier 3+) & Repair. * **Refactor:** **Decoupled**. * No longer called implicitly during the build loop (prevents network timeouts). * Invoked explicitly via `python manage.py generate`. * **Architect:** Generates raw GF code from scratch. * **Surgeon:** Patches broken `.gf` files based on compiler logs. ### C. The Scanners (Auditors) * **Scripts:** `rgl_auditor.py`, `lexicon_scanner.py` * **Role:** Quality Scoring. * **Logic:** analyze RGL coverage and Vocabulary depth to assign a **Maturity Score (0-10)**, enabling the Orchestrator to choose between "High Road" (Tier 1) and "Safe Mode" (Tier 3). --- ## 5. Level 4: The Runtime Services 🔌 Once Level 2 finishes successfully, Level 1 spawns these persistent processes in visible windows. ### A. The API (Brain) * **Script:** `app/adapters/api/main.py` * **Framework:** FastAPI + Uvicorn * **Role:** Handles HTTP requests, manages the Dependency Injection Container, and serves the Swagger UI. ### B. The Worker (Muscle) * **Script:** `app/workers/worker.py` * **Framework:** ARQ (Redis) * **Optimization:** **OS-Native Hot Reload**. * **Logic:** Uses `watchfiles` (instead of polling) to reload the `semantik_architect.pgf` binary into memory the *instant* the builder updates it. --- ## Data Flow Summary 1. **Filesystem** (Lexicon/RGL) ➔ **Indexer** (Cache Check) ➔ **Matrix JSON**. 2. **Matrix JSON** ➔ **Orchestrator** (Parallel Build) ➔ **Factory** (Tier 3 Gen) ➔ **GF Compiler** ➔ **PGF Binary**. 3. **PGF Binary** ➔ **Worker** (Hot Reload) ➔ **API** (User Request). ================================================================================================ FILE: docs/Technical-Reference/ADR 005 Dual-Path Input Validation (Strict_vs_Prototype).md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 17db741749771ac53c66d918a576006685cf95455e5283a2070144d76c4830e3 CONTENT_BYTES: 4024 ================================================================================================ # ADR 005: Dual-Path Input Validation (Strict vs. Prototype) ### 1. Context and Problem Statement The SemantiK Architect operates in two distinct phases of the software lifecycle: 1. **Architecting (Draft Mode):** Developers and AI agents must invent new grammar functions (e.g., `mkIsAProperty`, `mkRobot`) rapidly to test linguistic theories. 2. **Publishing (Production Mode):** The system generates high-reliability content using agreed-upon standards and shared libraries. **The Problem:** The system currently enforces **Ninai Protocol** validation globally. Ninai acts as a strict "bouncer," rejecting any function call that has not been pre-registered in its library. This creates a development bottleneck: to test a single line of new GF grammar, the developer must first fork, edit, and reinstall the Python validation library. This blocks the "Sketchy/Draft" phase of development. ### 2. The Decision We are adopting a **Dual-Path Architecture** for the Generation API (`/generate`). The system will effectively "relax" the API contract to accept two distinct input schemas: * **Path A: Strict Mode (The "Green" Path)** * **Schema:** `ninai.constructors.Statement` * **Behavior:** Validates inputs against the strict Ninai allowlist. * **Target Audience:** Production services, Regression tests, Team interoperability. * **Error Handling:** Fails *before* execution if the semantic frame is invalid. * **Path B: Prototype Mode (The "Red" Path)** * **Schema:** `UniversalNode` (Recursive Generic Object) * **Behavior:** Accepts *any* function name and argument list. It acts as a transparent pass-through to the GF Compiler. * **Target Audience:** The Architect Agent, Human Prototyping, AI Hallucinations. * **Error Handling:** Fails *during* execution (Runtime Error) if the GF compiler rejects the function. ### 3. Architecture Diagram **Flow Logic:** 1. **API Router:** Receives the JSON payload. 2. **Dispatcher:** * Attempts to parse as **Strict Ninai**. * If that fails (or if the structure explicitly matches the Generic schema), it falls back to **Prototype Mode**. 3. **Engine Adapter:** Uses "Duck Typing" (Dynamic Typing) to read the `.function` and `.args` attributes from *either* object type. 4. **GF Runtime:** Executes the command. ### 4. Technical Specification: The Universal Node To support the Prototype Path, we define a recursive schema that mimics the *structure* of a syntax tree without enforcing the *content*. **Schema Definition (JSON Semantic):** ```json { "function": "String (Required) - The name of the GF operation", "args": [ "String | Integer | Float", "OR Another UniversalNode Object (Recursive)" ] } ``` **Example Payload (Prototype):** ```json { "function": "mkIsAProperty", "args": [ "Sky", { "function": "mkColor", "args": ["Blue"] } ] } ``` ### 5. Consequences | **Positive** | **Negative** | | --- | --- | | **Velocity:** Developers can test new grammar functions immediately without touching Python code. | **Runtime Errors:** Typos (e.g., `mkBio` vs `mkBoi`) will not be caught until the GF binary actually tries to run, potentially causing obscure error messages from the C-runtime. | | **AI Compatibility:** The "Architect Agent" (LLM) can hallucinate new grammar structures that validly execute, allowing for autonomous grammar discovery. | **Documentation Drift:** Since functions don't need to be registered, the API documentation (Swagger) won't automatically list all available grammar commands. | | **Interoperability:** We maintain full compatibility with the Ninai standard for the wider team. | | ### 6. Implementation Strategy 1. **Domain Layer:** Define the `UniversalNode` Pydantic model. 2. **API Layer:** Update `routes.py` to use `Union[Statement, UniversalNode]` for the request body. 3. **Adapter Layer:** Refactor `gf_wrapper.py` to serialize both object types into GF Linearize commands. --- **Status:** `ACCEPTED` **Date:** 2025-12-23 **Author:** SemantiK Architect Team ================================================================================================ FILE: docs/Technical-Reference/ADR 006 Human-in-the-Loop Grammar Generation and GF Codex.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: a67e2a631589aed58eae8ebf4459881d8a4a4a12f697c618871f8009cb57b456 CONTENT_BYTES: 4611 ================================================================================================ # ADR 006: Human-in-the-Loop (HITL) Grammar Generation & GF Codex **Status:** Accepted **Date:** February 2026 **Context:** SemantiK Architect v2.5 ## 1. Context and Problem Statement In the v2.0 architecture, the Architect Agent was designed to autonomously generate grammars for under-resourced languages (Tier 3) at compile time (Runtime Auto-Generation). While functional in theory, this approach revealed critical limitations in production: * **Probabilistic Instability ("Fire and Pray"):** Generating Grammatical Framework (GF) code via LLMs often produces complex typing errors. The "Surgeon" agent sometimes loops endlessly without successfully patching the strict syntax errors enforced by the GF C-compiler. * **Cost and Latency (Financial Drain):** Repeated attempts (retry loops) with the Gemini API upon every compilation failure consume excessive credits and significantly slow down the build pipeline. * **Finite Scope:** The number of target languages for Wikipedia is not infinite (approximately 300 languages). Dynamic, "blind" automation on every build is disproportionate to the actual need, which is simply to generate a static file permanently. ## 2. The Decision We are abandoning fully autonomous runtime generation in favor of a **Human-in-the-Loop (HITL)** model. The AI is no longer an "autonomous builder" but a **copilot**, guided by a human operator using a strict instruction manual called the **"GF Codex"**. Grammars generated this way and validated by a human will now be permanently stored in `gf/contrib/{lang}/` (Tier 2 - Manual Overrides), rather than the ephemeral `gf/generated/src/` folder. ## 3. The "GF Codex" (RAG / Few-Shot Context Injection) To ensure high-quality code on the first attempt, requests to the LLM will no longer rely on a simple static prompt. Instead, they will include a comprehensive "GF Codex." This reference document will contain: 1. **Anti-Crash Rules (Implementation Rules):** * *The Inlining Rule:* Absolute prohibition of using `let` variables inside `lin` blocks to prevent the `variable #0 is out of scope` error. * *The Symbolic Rule:* Mandatory use of `symb` for raw strings instead of `mkPN` to prevent `unsupported token gluing` errors. 2. **Strict Skeletons:** Strict definition of the types imposed by `semantik_architect.gf` (e.g., mandatory use of `Predicate = VP ;` and an absolute ban on hallucinating `VPS`). 3. **"Few-Shot" Examples:** Perfect templates of validated GF grammars (e.g., one SVO model, one SOV model) to guide the AI's coding style. ## 4. The New Deployment Workflow 1. **Initialization:** The operator identifies a missing language and launches the interactive `ai_refiner` tool from the developer dashboard (`/tools`). 2. **Generation (Copilot):** The tool sends the "GF Codex" and the typological order (SVO, SOV) of the target language to the LLM API. 3. **Human Validation:** The operator receives the GF draft, reviews it, and uses `tools/language_health.py --mode compile` to ensure the code compiles perfectly with the RGL library. 4. **Save & Commit:** Once the file compiles successfully, it is manually pushed to the repository under `gf/contrib/{lang}/Wiki{Lang}.gf`. 5. **Deterministic Build:** During the regular pipeline (`build_300.py` / `orchestrator.py`), the orchestrator detects the file in `contrib/` and links it directly into the `semantik_architect.pgf` binary without making any API calls. ## 5. Consequences | Impact | Description | | --- | --- | | **Positive** | **Drastic API Cost Reduction:** LLM calls are only made once per language, entirely outside the daily build pipeline. | | **Positive** | **Build Stability:** The build pipeline becomes 100% deterministic again. Unexpected errors caused by AI hallucinations during compilation are eliminated. | | **Positive** | **Quality Increase:** Human review guarantees that the grammar makes linguistic sense before it reaches production. | | **Negative** | **Operational Friction:** Adding a new language now requires human intervention (this trade-off is deemed highly acceptable given the finite limit of ~300 languages). | --- ### 💡 Recommended Code Updates: 1. **In `05-AI_SERVICES.md**`: Update the section on "The Architect" to specify that it is now a tool invoked manually via the `/tools` interface, rather than automatically triggered by `orchestrator.py`. 2. **In `builder/orchestrator.py**`: Remove the AI fallback loop (The Surgeon) triggered when compilation fails. The orchestrator should simply skip the broken/missing language (`SKIP`) and log an alert that human intervention is required. ================================================================================================ FILE: docs/Technical-Reference/ADR-001-GF_VERSION_STRATEGY.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 47cb5f33ea2beaa5a92629758c38a5da5f70fb1edc50fd714eaeca480ff2a775 CONTENT_BYTES: 11016 ================================================================================================ # GF Architecture & Developer Guide **Version:** 2.4 **Last Updated:** 2026-02-20 **Context:** SemantiK Architect (Python + Grammatical Framework) ## 1. Architectural Overview This system bridges two fundamentally different paradigms: 1. **Python Domain Layer:** Dynamic, object-oriented, string-heavy. 2. **GF Engine Layer:** Statically typed, functional abstract syntax trees. The critical challenge is the **type mismatch**. Python sees `"Marie Curie"` as a string. GF sees a typed term (e.g., `NP`, `PN`, etc.). Our architecture uses a **Bridge Pattern**: raw runtime data is wrapped into safe GF types before linearization. ### The Stack | Layer | File/Component | Responsibility | | --- | --- | --- | | **Abstract** | `gf/semantik_architect.gf` | Defines the API contract (schema). | | **Concrete (App Grammars)** | `gf/Wiki{WikiCode}.gf` | Implements the schema using the RGL (one per language). | | **Bridge (Syntax Instances)** | `generated/src/Syntax{RglCode}.gf` | Provides `Syntax{RglCode}` instances used by app grammars. | | **Library** | GF Resource Grammar Library (RGL) | Linguistic primitives (`mkS`, `mkCl`, etc.). | | **Adapter** | `app/adapters/engines/gf_wrapper.py` | Converts Pydantic objects → GF trees (PGF expressions). | --- ## 2. Naming: ISO vs WikiCode vs RglCode (Source of Most Bugs) ### Canonical rule: naming is driven by `data/config/iso_to_wiki.json` Confirmed mappings in this repo: - `en` → `{"wiki": "Eng"}` → app grammar is `gf/WikiEng.gf` - `fr` → `{"wiki": "Fre"}` → app grammar is `gf/WikiFre.gf` **Implication:** this codebase does **not** use `WikiEn.gf` / `WikiFr.gf`. ISO-2 codes do not appear directly in grammar filenames. ### Terms - **ISO**: `en`, `fr`, `de` (ISO-639-1) - **WikiCode**: `Eng`, `Fre`, `Ger`, … (from `iso_to_wiki.json`) - **RglCode**: `Eng`, `Fre`, `Ger`, … (the GF module suffix that your bridge instance targets) **Important nuance:** WikiCode and RglCode often match, but the **source of truth is different**: - WikiCode: `iso_to_wiki.json` - RglCode: what your RGL folder actually exposes (e.g., `GrammarEng.gf`, `ParadigmsEng.gf`) --- ## 3. Directory Structure & File Hygiene The system historically used two “generated” roots: - `generated/src` (**canonical**) - `gf/generated/src` (**legacy mirror**) On Windows-mounted filesystems (`/mnt/c/...`), symlinks can be unreliable. The commander supports both and will **sync** between them, but you must treat **`generated/src` as canonical**. ### ✅ Allowed / Canonical - `gf/semantik_architect.gf` — abstract syntax - `gf/Wiki{WikiCode}.gf` — app concrete grammars (e.g., `gf/WikiEng.gf`, `gf/WikiFre.gf`) - `generated/src/Syntax{RglCode}.gf` — bridge “Syntax instances” (e.g., `generated/src/SyntaxEng.gf`) - `gf/semantik_architect.pgf` — compiled binary (usually git-ignored) - `data/config/iso_to_wiki.json` — authoritative mapping (ISO → WikiCode) ### ⚠️ Allowed but Legacy (should be unified) - `gf/generated/src/*` — legacy generated location Prefer to make it a symlink to `generated/` if your FS supports symlinks; otherwise keep it synchronized and do not hand-edit. ### ❌ Prohibited / Remove or Avoid Creating - `gf/Wiki.gf` — legacy/ambiguous - `gf/WikiEn.gf`, `gf/WikiFr.gf` — wrong naming convention for this repo - `gf/Symbolic*.gf` — **CRITICAL:** local files can conflict with the RGL’s `Symbolic` modules - Any `*.RGL_BROKEN` variants under include paths — can shadow correct modules depending on search path order **Rule of thumb:** if you see “Generated dirs distinct” warnings, assume stale/shadowing risk until you’ve re-synced and rebuilt. --- ## 4. GF + RGL Versioning (Alignment Contract) ### 4.1 Runtime toolchain reality - **GF Core (compiler/runtime):** The current upstream “Latest” GF core release is **GF 3.12**. :contentReference[oaicite:0]{index=0} - **RGL is separate:** Upstream packages treat **gf-rgl** as its own repo/artifact (not “bundled inside” GF core anymore). :contentReference[oaicite:1]{index=1} ### 4.2 This project’s contract - `builder/orchestrator.py` refuses to compile if `gf-rgl` is not pinned to the configured ref/commit. - `python manage.py align --force` is the canonical entrypoint to: 1) pin `gf-rgl` to the expected ref, and 2) regenerate Tier-1 bridge/app grammars. ### 4.3 Pinning: use a ref that exists in *your* gf-rgl clone Your repo has encountered failures because the pin ref/commit didn’t exist locally (even after fetch). The robust rule: - **Pin by “ref” (tag/branch/commit), not by a magic short hash.** - Default pin must be a ref that actually exists in `gf-rgl`. Upstream `gf-rgl` tags currently include (examples): `20250812`, `20250429`, `GF-3.10`, `RELEASE-3.9`, `RELEASE-3.8`. :contentReference[oaicite:2]{index=2} (And the tag page itself notes that the “latest tag is called release-3.12”, but that ref may not be present in all clones/forks; always pin to what *your* repo can resolve.) :contentReference[oaicite:3]{index=3} **Practical guidance:** - Treat `gf-rgl/` as a normal git clone (submodule configuration is **not assumed**). - Your alignment tool must: - `git -C gf-rgl fetch --tags` (or equivalent) - validate the ref exists - checkout/reset to it --- ## 5. The “Safe” RGL API (Reference) We restrict our RGL usage to a stable subset. ### Core Semantic Constructors | Function | Signature | Description | | --- | --- | --- | | `mkS` | `Cl -> S` | Clause → Sentence | | `mkCl` | `NP -> VP -> Cl` | Predication (“John walks”) | | `mkNP` | `Det -> N -> NP` | Determination (“the animal”) | | `mkVP` | `V2 -> NP -> VP` or `VP -> NP -> VP` | Transitive VP | | `mkAP` | `A -> AP` | Adjectival phrase | ### Structural Helpers | Function | Type | Usage | | --- | --- | --- | | `and_Conj` | `Conj` | List conjunction | | `in_Prep` | `Prep` | “in” | | `symb` | `String -> NP` | **Type bridge** for raw strings | > Note: the exact module providing `symb` depends on RGL version; do not shadow `Symbolic*` locally. --- ## 6. Implementation Rules (Anti-Crash / Stability Rules) ### Rule #1: Inlining Rule (Scope Safety) Avoid `let` inside `lin` rules when composing complex RGL macros. **Bad:** ```haskell mkEvent subj obj = let v = mkV "participate" in mkS (mkCl subj (mkVP v obj)) ```` **Good:** ```haskell mkEvent subj obj = mkS (mkCl subj (mkVP (mkV "participate") obj)) ``` ### Rule #2: Symbolic Rule (Runtime String Safety) Do not run morphology over runtime variables with `mkPN`/`mkN` when the input is unknown at compile time. * **Bad:** `mkLiteral s = mkNP (mkPN s)` * **Good:** `mkLiteral s = symb s` ### Rule #3: Modifier Type Rule If you define a `Modifier`, use `Adv` (not `AdV`) unless you have a controlled, language-specific reason. --- ## 7. The Python Adapter Pattern The Python wrapper (`gf_wrapper.py`) must construct ASTs using grammar bridge functions, not raw strings directly. ```python # Wrong: raw strings passed directly # pgf.Expr("mkBio", ["Marie", "Physicist"]) # Right: wrap strings using bridge constructors that exist in semantik_architect.gf subj = pgf.Expr("mkLiteral", [pgf.readExpr('"Marie"')]) prop = pgf.Expr("mkStrProperty", [pgf.readExpr('"Physicist"')]) expr = pgf.Expr("mkBio", [subj, prop]) ``` **Rule:** the adapter must match the exact function names + arities in `gf/semantik_architect.gf`. --- ## 8. Troubleshooting Dictionary | Symptom / Error | Diagnosis | Fix | | -------------------------------------------------------------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------------------------ | | `Cannot connect to the Docker daemon at unix:///var/run/docker.sock` (WSL) | WSL Linux docker socket isn’t wired to Docker Desktop | Enable Docker Desktop WSL integration *or* ensure your tooling uses `docker.exe` from WSL. | | `gf-rgl is not pinned to the expected ...` | RGL pin mismatch | Run alignment (or set the pin ref to something that exists locally). | | `Ref 'release-3.xx' not found in gf-rgl (even after fetch)` | The ref doesn’t exist in your clone/fork | Pick an existing tag in your `gf-rgl` (e.g., `20250812`) and pin to that. | | `fatal: No url found for submodule path 'gf-rgl' in .gitmodules` | Your repo is not using gf-rgl as a submodule | Alignment must treat `gf-rgl/` as a normal git clone; remove submodule assumptions. | | `SyntaxX.gf does not exist` | Missing bridge instance | Run `python tools/bootstrap_tier1.py --force` (or `manage.py align`). | | `atomic term conflict` / `Symbolic*` conflicts | Local `Symbolic*.gf` shadowing RGL | Delete local `gf/Symbolic*.gf`. | | `Function ... not found` | PGF stale or wrong grammar set compiled | Rebuild PGF and restart API. | | “Generated dirs distinct” warnings | `generated/src` and `gf/generated/src` diverged | Prefer `generated/src`; re-sync and rebuild; don’t hand-edit legacy mirror. | | `'venv/bin/python' is not recognized` (PowerShell) | `manage.py` assumes Unix venv layout | Run build commands in WSL, or update `manage.py` to resolve venv python per-OS. | --- ## 9. Build Commands (Current Canonical Flow) ### Recommended: build in WSL The default `manage.py` configuration assumes Unix venv paths (`venv/bin/python`). Running `manage.py build` in Windows PowerShell will fail unless you provide a Windows venv layout and/or a Windows-aware venv resolver. ### Minimal “Known Good” Flow (Tier-1, small set) ```bash # 1) Align (pins gf-rgl + generates Tier-1 bridges/app grammars) python manage.py align --force # 2) Build only a small language set while iterating python manage.py build --langs en fr ``` ### If you need to refresh RGL inventory for the matrix ```bash python tools/everything_matrix/build_index.py --regen-rgl python tools/bootstrap_tier1.py --force ``` ### Cleaning Do **not** delete all `gf/Wiki*.gf` anymore — those are canonical app grammars in this repo (WikiCode naming). Prefer: ```bash python manage.py clean ``` If you must do a manual clean, focus on compiled artifacts (`*.gfo`, `*.pgf`) and generated outputs, not source grammars. ================================================================================================ FILE: docs/Technical-Reference/CURRENT_RUNTIME_STATUS.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 12acfc3b826a52b806764685c8cbae65c5083af0645cb8e24d4323825afd79a3 CONTENT_BYTES: 11235 ================================================================================================ # CURRENT RUNTIME STATUS Last checked: 2026-03-19 ## Purpose This document records the **current observable runtime behavior** of SemantiK Architect **for the post-cutover EN/FR runtime state**. It is an operations/status page, not a target-state architecture spec. It does not redefine architecture, contracts, or acceptance. Its role is to describe the runtime state that is considered **current and true** once the EN/FR final cutover is in place. Where target architecture, runtime contract, public contract, cutover sequencing, and acceptance criteria all converge, this document records the **current live result** of that convergence. --- ## 1. Executive summary SemantiK Architect is currently in a **planner-first runtime state** for the EN/FR bio/person slice. In practical terms, this means: * the nominal single-sentence generation path is planner-first, * the canonical internal runtime handoff is `ConstructionPlan -> SurfaceResult`, * the public `/api/v1/generate/{lang_code}` route remains the stable generation entrypoint, * EN bio/person generation resolves to `WikiEng` and surfaces English, * FR bio/person generation resolves to `WikiFre` and surfaces French, * the public success envelope is stable and explicit, * and language correctness is evaluated by routing, runtime path, contract validity, and surface correctness together. Compatibility support may still exist at specific ingestion or fallback edges, but it is **not** the current runtime center of truth and it does **not** define nominal success. --- ## 2. Public HTTP surface currently mounted The backend is canonically served under `/api/v1/...`. Currently mounted routes include: * generation under `/api/v1/generate/...`, * health under both `/health/*` and `/api/v1/health/*`, * public language/entity/frame endpoints under `/api/v1/...`, * management endpoints under `/api/v1/...`, * tools under `/api/v1/tools/...`. Dual health mounting remains intentional so probes and API consumers can both use health routes without path-rewrite assumptions. --- ## 3. Current generation endpoint The primary generation route is: `POST /api/v1/generate/{lang_code}` This is the canonical route used by frontend/tooling, smoke checks, and runtime validation flows. ### Current public response shape Successful generation requests currently serialize to one canonical JSON envelope centered on: * `text` * `lang_code` * `construction_id` * `renderer_backend` * `fallback_used` * `tokens` * `debug_info` * `generation_time_ms` Current interpretation: * `text` is authoritative, * `lang_code` identifies the returned surface language, * `construction_id` is explicit on the nominal path, * `renderer_backend` is explicit on the nominal path, * `fallback_used` is explicit, * `tokens` correspond to the final surface text, * `generation_time_ms` is top-level and authoritative, * `debug_info` must not contradict top-level fields. Older success expectations centered on `surface_text` / `meta` are not aligned with the current public contract. --- ## 4. Request normalization currently supported ### 4.1 Language resolution `{lang_code}` in the URL is currently **authoritative**. If the payload also includes a language field (`lang`, `language`, `lang_code`, or `inputs.language`), it must normalize to the same language as the URL or the request is rejected. If the URL does not provide a language, the payload must provide one. Language normalization currently: * lowercases, * strips a leading `wiki...` prefix if present, * canonicalizes through shared lexicon code normalization. ### 4.2 Bio/person compatibility ingestion The runtime still accepts multiple legacy and compatibility aliases for biography/person generation at the request boundary, including: * `bio` * `biography` * `entity.person` * `entity_person` * `person` * `entity.person.v1` * `entity.person.v2` These inputs are normalized into the current bio/person frame path before generation. Current interpretation: * compatibility ingestion is permitted, * but compatibility ingestion does not redefine the nominal runtime, * and it does not weaken the planner-first acceptance gate. ### 4.3 Prototype / Ninai support If the incoming payload contains a top-level `function`, it is treated as a Ninai-style / prototype-style payload and routed through Ninai parsing rather than the standard frame parser. This support may still exist as a compatibility/prototype boundary, but it is not the canonical production semantics for the EN/FR final bio/person slice. --- ## 5. Current runtime center of truth ### 5.1 Current nominal runtime path The nominal runtime path is: `canonical input -> planner -> lexical resolution -> realizer -> SurfaceResult -> API response mapping` The canonical runtime handoff is: `ConstructionPlan -> SurfaceResult` This is the runtime truth for the current EN/FR bio/person slice. ### 5.2 Current runtime interpretation Current operational interpretation is: * planner-first is the nominal runtime, * direct legacy generation is not the current target-state path, * compatibility-only success does not count as nominal success, * runtime truth must exist before API mapping, * and the response mapper serializes canonical results rather than inventing nominal metadata for the first time. ### 5.3 Observable runtime metadata The runtime is expected to expose structured `debug_info`. On the nominal planner-first path, current runtime metadata is expected to make the following visible: * `runtime_path` * `fallback_used` * `renderer_backend` * `construction_id` * `lang_code` Additional structured metadata may include values such as: * `resolved_language` * `selected_backend` * `attempted_backends` * `slot_keys` * `backend_trace` * truthful compatibility markers when compatibility behavior is actually used Current interpretation: * `debug_info` mirrors runtime truth, * it does not replace required top-level public fields, * and compatibility metadata never upgrades a compatibility path into nominal success. --- ## 6. Health and validation endpoints The current runtime exposes: * `/health/live` * `/health/ready` * `/api/v1/health/live` * `/api/v1/health/ready` These endpoints remain part of the expected public runtime surface. The recommended validation path for runtime status remains: 1. refresh matrix, 2. validate lexicon, 3. compile PGF, 4. run runtime/language health, 5. run real EN and FR generation requests, 6. run relevant tests, 7. run `eval_bios`, 8. ensure docs reflect the achieved truth. --- ## 7. Current EN / FR status ### 7.1 English English bio/person generation is currently accepted for the target EN/FR cutover scope. Current operational interpretation: * `/api/v1/generate/en` is accepted, * the request resolves to `WikiEng`, * the nominal runtime path is planner-first, * `fallback_used = false` on nominal success, * the response contains the canonical public success envelope, * and the final surface text is English. ### 7.2 French French bio/person generation is currently accepted for the target EN/FR cutover scope. Current operational interpretation: * `/api/v1/generate/fr` is accepted, * the request resolves to `WikiFre`, * the nominal runtime path is planner-first, * `fallback_used = false` on nominal success, * the response contains the canonical public success envelope, * and the final surface text is French. Hard rule: * FR routed to `WikiFre` but still surfacing English is **not** a partial success, * it is a hard acceptance failure, * and it is not part of the current accepted runtime state. ### 7.3 EN / FR scope note The current accepted scope described here is: * EN bio/person generation * FR bio/person generation This document does not claim that all languages or all constructions have reached the same state. --- ## 8. What is stable right now The following are current runtime facts: * `/api/v1/generate/{lang_code}` is the canonical public generation route. * URL language is authoritative over payload language. * planner-first is the nominal runtime for the EN/FR bio/person slice. * the canonical internal runtime handoff is `ConstructionPlan -> SurfaceResult`. * EN resolves to `WikiEng`. * FR resolves to `WikiFre`. * FR success requires actual French surface output. * the public response contract is JSON with `text`, explicit structured runtime metadata, and authoritative top-level fields. * health endpoints are mounted in both root and `/api/v1` forms. * runtime validation still includes real generation calls as part of readiness/acceptance proof. --- ## 9. What remains intentionally compatibility-scoped The following may still exist at the boundary or compatibility edges without changing the current nominal runtime truth: * legacy bio/person request aliases, * Ninai/prototype-style input handling where still wired, * compatibility parsing of older internal result shapes, * compatibility/debug markers that truthfully describe fallback or migration behavior. Current interpretation: * these are compatibility surfaces, * not the architectural center of truth, * not the nominal planner-first success path, * and not an alternate acceptance model. --- ## 10. What is no longer correct to say The following are **not** correct current-state descriptions after the EN/FR final cutover: * that planner-first is only a target and not the current nominal runtime for EN/FR bio/person generation, * that direct frame-to-engine generation remains the primary EN/FR bio/person path, * that FR is merely routable but not validated for correct French surface output, * that `surface_text` / `meta` remain the canonical public success contract, * that routed-but-English FR output can still be counted as operational success, * that current runtime truth can be inferred from compatibility behavior alone. --- ## 11. Current status statement As of 2026-03-19, SemantiK Architect should be described as: > a working API with stable public generation and health endpoints, planner-first nominal runtime for the EN/FR bio/person slice, a canonical `ConstructionPlan -> SurfaceResult` internal contract, explicit top-level public success fields, English owned by `WikiEng`, French owned by `WikiFre`, and EN/FR acceptance that requires both contract correctness and language-correct surface realization. --- ## 12. Relationship to other documents This document is a **status page** only. It must be read consistently with: * `docs/architecture/multilingual_runtime_target.md` * `docs/contracts/construction_runtime_contract.md` * `docs/contracts/public_generation_response_contract.md` * `docs/contracts/public_vs_runtime_vs_frontend_boundaries.md` * `docs/contracts/debug_info_contract.md` * `docs/migration/en_fr_cutover_plan.md` * `docs/testing/EN_FR_bio_acceptance.md` * `docs/architecture/EN_FR_FINAL_PARALLEL_LOCKDOWN.md` Conflict rule: * architecture and contracts define what is authoritative, * the cutover plan defines completion sequencing, * `EN_FR_bio_acceptance.md` defines the operative EN/FR release gate, * the lockdown doc defines operational interpretation during the final parallel update, * this document only records the resulting current runtime truth. ================================================================================================ FILE: docs/Technical-Reference/DECISION_LOG.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: b70415d137d9fac9af90087c4e3119c3ca447575a79b91a674a1c423e9ea77ea CONTENT_BYTES: 10573 ================================================================================================ # DESIGN_DECISIONS.md SemantiK Architect – Design Decisions This document records the **key architectural choices** behind SemantiK Architect and the alternatives that were considered. It is meant to be a concise “why we did it this way” reference for reviewers and future contributors. --- ## 1. Overall System Shape ### Decision: Router → Family Engine → Constructions → Morphology → Lexicon **What we do** - The pipeline is: 1. **Router** (`router.py`) chooses: - family engine, - language profile, - morphology config, - lexicon. 2. **Semantics + Discourse** build a language-independent frame: - e.g. `BioFrame`, plus `DiscourseState`. 3. **Constructions** choose a clause pattern for the frame. 4. **Family Engine** realizes words + syntax using: - family matrix JSON, - language card JSON, - lexicon entries. 5. **Morphology** handles inflection, agreement, phonology details. 6. **Lexicon subsystem** provides lemma-level information and IDs. **Alternatives considered** 1. **Per-language renderers** (one function per language, no shared engine) - Pros: - Simple to start with. - Easy to hack per-language behavior. - Cons: - Massive code duplication. - Hard to keep consistent across 100+ languages. - Any bug fix must be copied by hand many times. 2. **Single monolithic “universal” renderer** - Pros: - Centralized logic, no engine routing. - Cons: - Would quickly become unmaintainable. - Hard to reason about language-specific code paths. - Difficult for contributors to work in a huge file. **Why we chose this** - We want **maximum reuse** across languages with **clear modularity**. - The pipeline matches how many NLG / grammar engineering systems are structured: - semantics → constructions → morphology/lexicon. - It allows **family engines** to share heavy logic, while still giving languages a way to override via data. --- ## 2. Family Engines Instead of Language-Specific Engines ### Decision: ~15 family engines (Romance, Germanic, Slavic, Agglutinative, Bantu, etc.) **What we do** - Implement a small number of **family-level engines** in `engines/*.py`. - Each engine is responsible for: - common morphosyntactic patterns for that family, - reading a shared family matrix JSON, - applying language-specific overrides via cards. **Alternative** - One engine per language (`engines/it.py`, `engines/es.py`, `engines/sw.py`, …). **Why we chose this** - Many languages share core behavior: - Romance: gender, articles, similar suffix morphology. - Slavic: case system, gendered past, similar declensional logic. - Bantu: noun class + concord across the clause. - Factoring this at the family level: - **reduces code duplication**, - makes it easier to add new languages in that family, - encourages us to think in **typological terms** (which is how AW often reasons). - Language-specific irregularities are encoded in: - JSON cards, - lexicon entries, - small overrides, not separate engines. --- ## 3. Constructions as a Separate Layer ### Decision: `constructions/*.py` are language-agnostic sentence patterns **What we do** - Constructions know about: - argument structure (subject, object, possessor, etc.), - information structure (topic/focus), - clause-level word order and connectors (copula, complementizers, etc.). - They do **not** know how to inflect; they call the engine’s Morphology API. **Alternative** - Hard-code constructions inside each engine (or even each language). **Why we chose this** - Many constructions are cross-linguistic: - “X is a Y”, “X has Y”, “There is Y in X”, relative clauses, topic-comment. - Having an explicit constructions layer: - Allows us to **reuse constructions** across families. - Makes the system easier to reason about for linguists (they can see “the set of constructions”). - Separates sentence patterns from morphophonology and low-level syntax. - It also lines up well with: - **construction grammar** and **frame semantics** ideas, - where constructions are first-class objects. --- ## 4. Data-Driven Morphology and Configuration ### Decision: Morphology and configuration live in JSON matrices + cards **What we do** - Store morphosyntactic rules in JSON: - Family matrices in `data/morphology_configs/*.json`. - Per-language cards in `data//.json`. - Engines read these and never hard-code suffixes, article maps, noun classes, etc., unless strictly necessary. **Alternative** - Store all morphology directly in Python code (if/else trees, hard-coded tables). **Why we chose this** - JSON is: - editable by non-programmers, - easier to version and diff for rules, - suitable for Wikifunctions / Z-data style representations. - It enables the **“crowdsource the cards”** strategy: - A contributor can improve `ca.json` or `sw.json` without writing Python. - It provides a clear path for: - exporting / importing configs to and from AW / Wikidata objects, - building UI tools for editing language behavior. --- ## 5. Separate Lexicon Subsystem ### Decision: Lexicon is its own package (`lexicon/*`) and data (`data/lexicon/*.json`) **What we do** - Lexicon functionality is not hidden in engines. - We have: - `lexicon/types.py` – Lexeme and related structures, - `lexicon/loader.py`, `lexicon/index.py` – loading + indexing, - `lexicon/wikidata_bridge.py`, `lexicon/aw_lexeme_bridge.py` – integration with Wikidata/AW lexemes, - `data/lexicon/*.json` – actual lexicon files per language. **Alternatives** 1. Lexicon integrated into engines (per-language dictionaries inside Python). 2. Use only Wikidata at runtime with no local lexicon. **Why we chose this** - Separating lexicon: - makes it reusable by **multiple engines and constructions**, - allows **large-scale** lexicon management (coverage reports, schemas), - makes integration with Wikidata lexemes explicit and testable. - Local lexica: - avoid being strictly dependent on live network calls to Wikidata, - allow us to guarantee availability and performance, - can be built / refreshed offline from dumps. --- ## 6. Semantics and Discourse: Light but Real ### Decision: Minimal but explicit semantics (`semantics/*`) and discourse (`discourse/*`) **What we do** - Use small dataclasses like `BioFrame`, `Entity`, `Event`, `TimeSpan` for input. - Maintain a `DiscourseState` for: - tracking mentioned entities, - salience and last mention, - topic selection. - Provide modules for: - information structure (`discourse/info_structure.py`), - referring expressions (`discourse/referring_expression.py`), - simple discourse planning (`discourse/planner.py`). **Alternative** - Treat each sentence independently; no discourse model. - Encode semantics as ad-hoc dicts with no consistent structure. **Why we chose this** - Multi-sentence texts (biography leads, short descriptions) **need** some discourse model: - pronouns vs full names, - topic markers vs canonical word order, - ordering of information. - A fully formal semantic/discourse system (UMR, full AMR, etc.) would be heavy for the initial implementation. - This “light but real” layer: - is enough to demonstrate **non-trivial discourse competence**, - stays simple enough for contributors to understand, - provides a bridge to more formal semantic inputs from AW. --- ## 7. Test-First and QA-Focused Design ### Decision: QA tools and test suites are first-class parts of the architecture **What we do** - `qa/test_runner.py` and `qa_tools/universal_test_runner.py` are the default way to validate engines. - `qa_tools/test_suite_generator.py` helps create large CSV-based suites (often with LLM assistance). - `qa_tools/lexicon_coverage_report.py` measures lexicon coverage against tests. - Dedicated tests for lexicon loading and Wikidata bridges. **Alternative** - Rely primarily on manual testing and ad-hoc scripts. - Only add tests at the very end, language by language. **Why we chose this** - Abstract Wikipedia needs **trustworthy** renderers: - changes can affect many languages at once. - A structured QA pipeline: - catches regressions when modifying family matrices, - quantifies coverage and quality per language, - makes it easier to onboard new languages with confidence. - CSV-based suites are: - easy to inspect, - easy to generate semi-automatically (e.g. via LLM prompts), - easy to extend by the community. --- ## 8. Implementation Language and Style ### Decision: Plain Python + JSON, with modest typing **What we do** - Use Python for: - engines, constructions, semantics, discourse, lexicon logic. - Use JSON for: - morphology configs, - language cards, - lexicon files, - some datasets. - Use simple type hints and dataclasses where helpful. **Alternatives** - A fully typed/compiled language (e.g. Haskell, OCaml, Rust). - A DSL / custom language for grammars. **Why we chose this** - Python is: - accessible to many contributors, - already used in Wikimedia / AW-related tooling, - easy to prototype in and profile. - JSON is: - compatible with Wikifunctions / Z-data style, - familiar to Wikimedia/MediaWiki ecosystem, - easy to manipulate from many languages and tools. --- ## 9. Scope and Non-Goals ### Explicit non-goals (for now) - **Full coverage of all sentence types** in every language. - **Complete semantic theory** or full alignment with any one formalism (UMR, AMR, Ninai). - **Parsing**; the system is generation-focused. - **End-to-end styling / register control**, beyond what basic discourse choices allow. The design aims to: - Solve the **architecture problem** for multilingual NLG in AW, - Provide a serious, extensible base that can be: - critiqued, - improved, - specialized for particular language families. --- ## 10. Summary The key design choices are: - **Family engines** instead of per-language engines. - A separate, **language-agnostic constructions layer**. - **Data-driven morphology** via family matrices and language cards. - A **distinct lexicon subsystem** integrated with Wikidata. - A **light but explicit semantics + discourse layer**. - A **test-first, QA-heavy** workflow. Together, these choices are intended to make SemantiK Architect: - scalable across many languages, - understandable by both engineers and linguists, - and robust enough for real-world Abstract Wikipedia use. ================================================================================================ FILE: docs/Technical-Reference/ENGINE_INTERNALS.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 5a0155fe6407fc56c1f0f6051bd8a64a5a564f68ec9ceda08a82e020dd246c6a CONTENT_BYTES: 5763 ================================================================================================ Here is the formal documentation for the **SemantiK Architect: Multilingual Engine V2**. This document defines the architecture, directory structure, and logic required to scale the system from \~40 to 300+ languages using a **Hybrid Factory** approach. ----- # 📘 Project Abstract: The Hybrid Multilingual Engine ### 1\. Core Philosophy To support the 300+ languages required by Abstract Wikipedia, we cannot rely solely on the academic **Resource Grammar Library (RGL)**, which covers only \~40 languages. Conversely, manual implementation of 260+ languages is unscalable. We adopt a **Three-Tier Hybrid Architecture**: 1. **Prioritize Quality:** Use official, expert-written grammars where available. 2. **Allow Overrides:** Enable manual community contributions for specific languages. 3. **Guarantee Coverage:** Programmatically generate "Pidgin" (simplified) grammars for all remaining languages to ensure 100% API availability. ### 2\. The "Waterfall" Lookup Logic The build system (`build_300.py`) will resolve a language code (e.g., `zul`) by checking sources in a specific priority order. * **Priority 1: Official RGL.** *Is it in the standard library?* If yes, use it. * **Priority 2: Contrib.** *Do we have a manual draft in `gf/contrib`?* If yes, use it (overrides RGL). * **Priority 3: Factory.** *Is it in the generated folder?* If yes, use it. * **Fail:** If none exist, skip the language. ----- ### 3\. File Arborescence (Directory Structure) The project structure is reorganized to separate **Source** (Official), **Manual** (Community), and **Generated** (Factory) assets. ```text C:\MyCode\SemantiK_Architect\Semantik_architect\ │ ├── architect_http_api\ │ └── gf\ │ └── language_map.py # [ROUTER] Maps Z-IDs (Z1002) to Concrete Names (WikiEng) │ ├── docker\ │ └── Dockerfile.backend # [ENV] Copies all 3 grammar folders into the container │ ├── gf\ │ ├── Wiki.pgf # [ARTIFACT] The compiled binary containing ALL languages │ ├── build_300.py # [BUILDER] Orchestrates the "Waterfall" compilation │ │ │ ├── gf-rgl\ # [TIER 1] Official RGL (Cloned Git Repo) │ │ └── src\ # Contains: english, french, chinese... │ │ │ ├── contrib\ # [TIER 2] Manual Overrides (Version Controlled) │ │ └── que\ # Example: High-quality manual Quechua grammar │ │ └── WikiQue.gf │ │ │ └── generated\ # [TIER 3] Factory Output (Git Ignored / Auto-Cleaned) │ ├── zul\ # Example: Auto-generated Zulu stubs │ ├── yor\ # Example: Auto-generated Yoruba stubs │ └── ... (250+ others) │ ├── utils\ │ └── grammar_factory.py # [GENERATOR] Reads config, writes files to 'gf/generated/' │ └── test_gf_dynamic.py # [VERIFIER] Dynamically tests every language in Wiki.pgf ``` ----- ### 4\. Component Definitions #### A. The Language Factory (`utils/grammar_factory.py`) * **Purpose:** To ensure no language is left behind. * **Input:** A configuration dictionary defining the "DNA" of missing languages (Name, ISO Code, Word Order). * **Output:** Valid, compilable GF source files (`Res`, `Syntax`, `Wiki`) implementing a simplified "Pidgin" grammar (e.g., SVO string concatenation). * **Lifecycle:** Runs *before* the build. Wipes and recreates the `gf/generated/` folder every time. #### B. The Orchestrator (`gf/build_300.py`) * **Purpose:** To bind all separate grammar files into a single Portable Grammar Format (`.pgf`) file. * **Logic:** * Iterates through a master list of 300+ ISO codes. * Resolves the file path for each code using the **Waterfall Logic**. * Injects the `gf/generated` and `gf/contrib` paths into the compiler's search scope. * Generates the bridging `Wiki{Lang}.gf` file for RGL languages to connect them to our API. #### C. The Router (`language_map.py`) * **Purpose:** To translate external identifiers into internal GF concrete grammar names. * **Logic:** * `get_concrete("fra")` -\> `WikiFre` (Legacy RGL naming) * `get_concrete("zul")` -\> `WikiZul` (Standard Factory naming) * `get_concrete("Z1002")` -\> `WikiEng` (Wikidata ID mapping) ----- ### 5\. Developer Workflows #### Scenario A: Adding a "Missing" Language * **Goal:** Add *Hausa* (`hau`), which is not in the RGL. * **Action:** Open `utils/grammar_factory.py` and add to the config: ```python "hau": {"name": "Hausa", "order": "SVO"} ``` * **Result:** The next build automatically creates `WikiHau` and compiles it. #### Scenario B: "Graduating" a Language * **Goal:** Replace the "Pidgin" Zulu grammar with a high-quality manual one. * **Action:** 1. Create folder `gf/contrib/zul/`. 2. Write (or paste) the high-quality `WikiZul.gf` files there. * **Result:** `build_300.py` detects the folder in `contrib`, ignores the one in `generated`, and compiles the high-quality version. #### Scenario C: Updating the Core * **Goal:** Get the latest fixes for English or French. * **Action:** Run `git pull` inside `gf/gf-rgl/`. ----- ### 6\. Implementation Checklist 1. ✅ **Architecture Defined.** 2. ⬜ **Create Factory Script:** `utils/grammar_factory.py`. 3. ⬜ **Update Orchestrator:** `gf/build_300.py` (Waterfall logic). 4. ⬜ **Update Router:** `language_map.py` (300+ code support). 5. ⬜ **Update Docker:** Copy `contrib` and `generated` folders. 6. ⬜ **Verification:** Run `test_gf_dynamic.py`. ================================================================================================ FILE: docs/Technical-Reference/everything_matrix_orchestration.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: ec1537c6b48f048e05d0602d9c9acec2e68183d6bb2ab034797f045e9cdbcfce CONTENT_BYTES: 6890 ================================================================================================ # Everything Matrix Orchestration This document describes the **single-orchestrator** architecture for the Everything Matrix suite. ## Goal `tools/everything_matrix/build_index.py` is the **only normal entrypoint** that refreshes the full matrix: * **Zone A**: RGL inventory + grammar completion signals (from `rgl_inventory.json`) * **Zone B**: Lexicon health * **Zone C**: App readiness * **Zone D**: QA readiness All other scripts in `tools/everything_matrix/` remain available as **debug tools**, and must be: * **side-effect free by default** * **not duplicate work** during a normal matrix refresh --- ## Outputs ### Primary output * `data/indices/everything_matrix.json` ### Cache / fingerprint output * `data/indices/filesystem.checksum` (or `matrix.output_dir/filesystem.checksum` if configured) Used to skip work when inputs are unchanged. ### Prerequisite artifacts * `data/indices/rgl_inventory.json` (Zone A source-of-truth input) --- ## Configuration ### Config file * `data/config/everything_matrix_config.json` ### Paths controlled by config (typical keys) * `matrix.output_dir` (default `data/indices`) * `matrix.everything_index` (default `data/indices/everything_matrix.json`) * `rgl.inventory_file` (default `data/indices/rgl_inventory.json`) * `rgl.src_root` (default `gf-rgl/src`) * `lexicon.lexicon_root` (default `data/lexicon`) * `qa.gf_root` (default `gf`) * `frontend.assets_path`, `frontend.profiles_path` * `backend.profiles_path` * `iso_map_file` (**should point at** `data/config/iso_to_wiki.json`) --- ## Normalization rules ### Canonical language key * The matrix is keyed by **ISO-639-1 `iso2`**, **lowercase** (example: `en`, `fr`, `sw`). ### No mixing of key types in orchestrator * `build_index.py` stores matrix entries under **iso2 only** * Any wiki/iso3 keys are normalized at scanner boundaries (preferred) or centrally before synthesis ### Source of truth for normalization `tools/everything_matrix/norm.py` is the shared module for: * loading `iso_to_wiki.json` * mapping `wiki` / `iso3` / `WikiXxx` forms → canonical `iso2` * mapping language display names **Canonical iso map location** is `data/config/iso_to_wiki.json`. A legacy `config/iso_to_wiki.json` may exist, but new code/config should point to the `data/config/` location. --- ## Scanner contracts Build Index uses scanners as libraries. The orchestrator calls each scanner **once per zone** (one-shot scan), then performs **dict lookups** per language. ### Zone A — RGL **Library contract** * `rgl_scanner.scan_rgl(...) -> inventory_dict` **Normal behavior** * `build_index.py` reads `data/indices/rgl_inventory.json` * It does **not** rescan `gf-rgl/src` unless `--regen-rgl` or inventory missing **Side effects** * `rgl_scanner.scan_rgl(write_output=False)` must be side-effect free * CLI can write with `--write` (debug only) ### Zone B — Lexicon **Library contract** * `lexicon_scanner.scan_all_lexicons(lex_root: Path) -> dict[, zone_b_stats]` **Expected keys** * `{"SEED": float, "CONC": float, "WIDE": float, "SEM": float}` **Scale** * all values are `0..10` floats ### Zone C — App readiness **Library contract** * `app_scanner.scan_all_apps(repo_root: Path) -> dict[, zone_c_stats]` **Expected keys** * `{"PROF": float, "ASST": float, "ROUT": float}` **Scale** * all values are `0..10` (ints/floats allowed; orchestrator clamps) ### Zone D — QA readiness **Library contract** * `qa_scanner.scan_all_artifacts(gf_root: Path) -> dict[, zone_d_stats]` **Expected keys** * `{"BIN": float, "TEST": float}` **Scale** * all values are `0..10` floats > Note: Scanners may emit keys as iso2/wiki/iso3 forms; `build_index.py` normalizes keys to iso2 before writing the matrix. --- ## Orchestrator behavior `tools/everything_matrix/build_index.py` performs: 1. **fingerprint/cache check** (skips work on cache hit unless forced) 2. ensures prerequisite inventories exist (optionally regenerates) 3. runs **one-shot scans** for each zone 4. synthesizes per-language verdict + maturity 5. writes `data/indices/everything_matrix.json` 6. writes the new fingerprint to `filesystem.checksum` ### Cache hit semantics * Default: if fingerprint matches, **returns early** * With `--touch-timestamp`: rewrites only the matrix timestamp fields on cache hit * With `--force` (or any `--regen-*`): bypasses cache ### One-shot scan rule During a normal refresh, build_index calls: * `lexicon_scanner.scan_all_lexicons(...)` **once** * `app_scanner.scan_all_apps(...)` **once** * `qa_scanner.scan_all_artifacts(...)` **once** Inside the per-language loop it does **only dict lookups**. ### No duplicate RGL scans During a normal refresh, build_index: * does **not** call `rgl_scanner.scan_rgl()` unless explicitly requested (`--regen-rgl`) or the inventory file is missing --- ## Scoring rules ### Scale All zone values are clamped to `0..10`. ### Back-compat shim `build_index.py` may rescale legacy `0..1` values into `0..10`, but scanners should emit `0..10`. ### Zone averages Per-language averages are computed as mean of each zone's sub-blocks: * `A_RGL` average of `CAT, NOUN, PARA, GRAM, SYN` * `B_LEX` average of `SEED, CONC, WIDE, SEM` * `C_APP` average of `PROF, ASST, ROUT` * `D_QA` average of `BIN, TEST` ### Maturity Maturity is a weighted sum of the zone averages: * weights live in `data/config/everything_matrix_config.json` (under `matrix.zone_weights` or equivalent) * output is clamped to `0..10` ### Strategy ladder The orchestrator produces one of: * `HIGH_ROAD` * `SAFE_MODE` * `SKIP` (Exact thresholds are configured in the matrix config.) --- ## Downstream consumer: Grammar build orchestrator The Everything Matrix **does not compile grammars**. It produces per-language `build_strategy` (and supporting signals) that downstream build tooling consumes. Canonical grammar build entrypoints: * **Programmatic**: `builder.orchestrator.build_pgf(...)` * **CLI**: `python -m builder.orchestrator ...` (preferred) * **Legacy shim** (if present): `gf/build_orchestrator.py` (wrapper only) --- ## CLI usage ### Normal run ```bash python tools/everything_matrix/build_index.py ``` ### Force rebuild (bypass cache) ```bash python tools/everything_matrix/build_index.py --force ``` ### Regenerate only the prerequisite RGL inventory (and rebuild) ```bash python tools/everything_matrix/build_index.py --regen-rgl ``` ### Force-rescan specific zones (and rebuild) ```bash python tools/everything_matrix/build_index.py --regen-lex python tools/everything_matrix/build_index.py --regen-app python tools/everything_matrix/build_index.py --regen-qa ``` ### Cache hit but refresh timestamps ```bash python tools/everything_matrix/build_index.py --touch-timestamp ``` ### Verbose logging ```bash python tools/everything_matrix/build_index.py --verbose ``` ================================================================================================ FILE: docs/Technical-Reference/everythingMatrixUpgrade.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 34bccb6bc842f1e1ff50e0937a58cbbf3126a453e2359706cce3a82b8a92b532 CONTENT_BYTES: 5920 ================================================================================================ # 🧠 The Everything Matrix v2.1: Health & Decision Specification **SemantiK Architect** ## 1. Executive Summary The **Everything Matrix** (`data/indices/everything_matrix.json`) is the central nervous system of the build pipeline. It replaces static configuration with a dynamic **Health Ledger**. In **v2.1**, the Matrix evolves from a simple "File Exists" checker to a **14-Point Deep Tissue Scanner**. It audits every language across four strategic zones (Logic, Data, Application, Quality) to calculate a precise **Maturity Score (0-10)**. This score empowers the Orchestrator to make autonomous decisions (e.g., *"French is robust enough for High Road compilation, but Zulu needs Safe Mode and AI-repair"*). --- ## 2. The 4 Zones & 14 Health Blocks Every language is graded on 14 specific metrics. These are not boolean flags; they are continuous scores (0.0 - 10.0) derived from physical code analysis. ### Zone A: RGL Engine (Logic) ⚙️ *Audits the structural integrity of the Grammatical Framework implementation.* | Block | Metric Name | Definition | Target (10/10) | | --- | --- | --- | --- | | **1** | `CAT` | **Categories** | Are standard types (`N`, `V`, `A`) defined? | | **2** | `NOUN` | **Morphology** | Can nouns/verbs be inflected? | | **3** | `PARA` | **Paradigms** | Are smart constructors (`mkN`, `mkV`) avail? | | **4** | `GRAM` | **Grammar** | Is the structural core implemented? | | **5** | `SYN` | **Syntax** | Is the API layer exposed? | > **Impact:** If `CAT` or `SYN` < 10, the build strategy automatically downgrades to **SAFE_MODE** (Factory) because the RGL is incomplete. ### Zone B: Lexicon (Data) 📚 *Audits the vocabulary depth and semantic alignment.* | Block | Metric Name | Definition | Target (10/10) | | --- | --- | --- | --- | | **6** | `SEED` | **Core Seed** | Size of functional vocabulary (`core.json`). | | **7** | `CONC` | **Concepts** | Size of domain vocabulary (`people.json`). | | **8** | `WIDE` | **Wide Import** | Existence of bulk-import CSV. | | **9** | `SEM` | **Semantics** | Wikidata Alignment Score. | > **Impact:** If `SEED` < 2.0, the language is marked `runnable: false` to prevent the runtime from crashing on empty dictionaries. ### Zone C: Application (Use Case) 🚀 *Determines readiness for specific vertical capabilities.* | Block | Metric Name | Definition | Requirement | | --- | --- | --- | --- | | **10** | `PROF` | **Bio-Ready** | Can generate Biographies? | | **11** | `ASST` | **Chat-Ready** | Can handle dialog? | | **12** | `ROUT` | **Routing** | Is topology configured? | ### Zone D: Quality (Verification) 🛡️ *Tracks the physical artifacts and regression status.* | Block | Metric Name | Definition | Requirement | | --- | --- | --- | --- | | **13** | `BIN` | **Binary** | Is present in `semantik_architect.pgf`? | | **14** | `TEST` | **Regression** | Gold Standard Pass Rate. | --- ## 3. The JSON Schema The `build_index.py` script aggregates all scanners into this finalized structure. ```json "fra": { "meta": { "iso": "fra", "name": "French", "family": "Romance" }, "zones": { "A_RGL": { "CAT": 10, "NOUN": 10, "PARA": 10, "GRAM": 10, "SYN": 10 }, "B_LEX": { "SEED": 8.5, "CONC": 4.2, "WIDE": 10, "SEM": 9.0 }, "C_APP": { "PROF": 1.0, "ASST": 0.0, "ROUT": 1.0 }, "D_QA": { "BIN": 1.0, "TEST": 0.8 } }, "verdict": { "maturity_score": 8.9, // Weighted Average (A*0.6 + B*0.4) "build_strategy": "HIGH_ROAD", // Decisions: HIGH_ROAD | SAFE_MODE | SKIP "runnable": true // If False, Worker will refuse to load it. } } ``` --- ## 4. The Scanning Architecture The Matrix is populated by three specialized "Census Takers" running in parallel. ### 1. `rgl_auditor.py` (Zone A) * **Target:** `gf-rgl/src/{lang}/` * **Logic:** Checks for the physical existence of the 5 standard GF modules. * **Output:** The structural integrity score. ### 2. `lexicon_scanner.py` (Zones B & C) * **Target:** `data/lexicon/{lang}/` * **Logic:** * Parses JSON shards to count entries (`SEED`, `CONC`). * Checks specific keys (`qid`, `forms`) for semantic alignment (`SEM`). * Validates domain readiness (e.g., checks if `people.json` contains "physicist" for `PROF` score). ### 3. `qa_scanner.py` (Zone D) * **Target:** `gf/` and `tests/logs/` * **Logic:** * Verifies if the language key exists in `semantik_architect.pgf`. * Parses the latest JUnit XML report from `pytest` to extract pass/fail rates. --- ## 5. Decision Logic (The Verdict) The Orchestrator (`builder/orchestrator.py`) reads the `verdict` object to determine the build path. ### Strategy Table | Maturity | Zone A Score | Verdict | Orchestrator Action | | --- | --- | --- | --- | | **> 7.0** | **10 (Perfect)** | `HIGH_ROAD` | Links directly to RGL source. Full optimization. | | **> 2.0** | **Any** | `SAFE_MODE` | Generates "Factory Grammar" (Weighted Topology). Triggers **Architect Agent** if file missing. | | **< 2.0** | **Any** | `SKIP` | Excludes language from build. | ### Runnable Logic The Worker (`app/workers/worker.py`) checks `verdict.runnable` before initialization. * **Rule:** `runnable = (SEED >= 2.0) OR (build_strategy == "HIGH_ROAD")` * **Why:** A language with no core words ("is", "the") will generate empty strings or crash the linearization engine. We protect the runtime by isolating these "Zombie Languages." --- ## 6. Implementation Guide ### Adding a New Language 1. **Register:** Create `data/lexicon/{iso}/core.json`. 2. **Scan:** Run `python tools/everything_matrix/build_index.py`. 3. **Check:** Ensure `SEED` score > 2.0. 4. **Build:** Run `python manage.py build`. ### Debugging Low Scores * **Low `SEM`:** Your JSON is missing `qid` fields. Run `harvest_lexicon.py`. * **Low `PROF`:** You cannot generate biographies. Add `people.json`. * **Low `TEST`:** Your grammar logic is flawed. Run `pytest` and check the **Judge Agent's** critique. ================================================================================================ FILE: docs/Technical-Reference/GF WordNet Structural Map.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 7cd5997caba0a642e82f6914a66261e3409b96cc51de6c72dad52d4c18b1916a CONTENT_BYTES: 4120 ================================================================================================ # 🗺️ GF WordNet Structural Map (v1.0) **Target Repository:** `gf-wordnet` **Purpose:** Lexicon Harvesting & Data Injection for SemantiK Architect. ## 1. The Core Architecture (The "Rosetta Stone") The repository is built on a **Star Topology**. The `WordNet` module is the center, and all 30+ languages are spokes that implement it. | Component | File Pattern | Role | Key Characteristic | | --- | --- | --- | --- | | **Abstract Interface** | `gf/WordNet.gf` | **Primary Key Database** | Defines function names (`apple_N`) and maps them to WordNet IDs (`02756049-n`). | | **Concrete Lexicon** | `gf/WordNet{Lang}.gf` | **Value Store** | Contains the actual words. E.g., `lin apple_N = mkN "pomme"`. | | **Assembler** | `Parse{Lang}.gf` | **Compiler Entry** | Binds the Lexicon (`WordNetFre`) to the Grammar (`ParseExtendFre`). | | **Extension** | `ParseExtend{Lang}.gf` | **Grammar Patch** | Adds extra grammatical rules (e.g., `CompVP`, `InOrderToVP`) not found in standard RGL. | --- ## 2. Directory Structure & Key Files Based on the file scan, here is the functional breakdown of the directory tree: ### 📂 Root / `gf/` * **`WordNet.gf`**: **CRITICAL.** The source of truth for semantic IDs. * *Format:* `fun function_name : Cat ; -- SynsetID`. * **`WordNet{Lang}.gf`** (e.g., `WordNetRus.gf`): **CRITICAL.** The target for harvesting. * *Format:* `lin function_name = constructor "word" ;`. * **`Parse.gf`**: The Abstract Grammar definition that enforces the `WordNet` dependency. * **`Parse{Lang}.gf`** (e.g., `ParseEng.gf`): The top-level concrete grammar. Useful for testing but **not** for harvesting words directly. ### 📂 `bootstrap/` * Contains Haskell scripts (`bootstrap.hs`, `build.hs`) used to generate the initial GF files from the Princeton WordNet database. * *Relevance:* Low for harvesting, High for understanding provenance. ### 📂 `morphodicts/` * **`MorphoDict{Lang}.gf`**: Contains raw inflection tables for complex languages (Arabic, etc.). * *Relevance:* **High.** If the harvester sees `variants {}` in `WordNet{Lang}.gf`, the word might be hidden here or missing. ### 📂 `www/` & `www-services/` * Web interface code for the Cloud GF WordNet browser. * *Relevance:* None for the build pipeline. --- ## 3. The Data Linking Protocol To extract a usable dictionary (`JSON Shard`), the harvester must traverse this specific path: ### Step 1: Extract the Semantic Key (Abstract) **Source:** `gf/WordNet.gf` **Regex:** `fun\s+(\w+)\s*:\s*\w+\s*;\s*--\s*([Q\d]+-?[a-z]*)` **Example Match:** * Function: `a_bomb_N` * ID: `02756049-n` ### Step 2: Extract the Lexical Value (Concrete) **Source:** `gf/WordNet{Lang}.gf` **Regex:** `lin\s+(\w+)\s*=\s*(.*?)\s*;` **Example Match (Rus):** * Function: `a_bomb_N` * RHS: `compoundN (mkA "атомный" "1*a") (mkN "бомба" feminine inanimate "1a")` * **Extracted Lemma:** "атомный", "бомба" (Heuristic: grab string literals). ### Step 3: Semantic Alignment (Wikidata) The `WordNet.gf` file contains two types of IDs in the comments: 1. **WordNet IDs:** `02756049-n` (8 digits + pos tag). 2. **Wikidata QIDs:** `Q25287` (e.g., `fun gothenburg_1_LN : LN ; -- Q25287`). **Action:** The harvester must detect `Q` prefixes and save them to the `qid` field in the JSON output. This enables the **"Q42 Killer"** logic in the Architect. --- ## 4. Known "Gotchas" & Edge Cases 1. **`variants {}`**: * Many definitions in `WordNetRus.gf` use `variants {}`. * *Meaning:* The word is **missing** in that language. * *Action:* The harvester must skip these entries to avoid polluting the database with empty strings. 2. **`--guessed`**: * Comments like `--guessed` appear frequently in `WordNetRus.gf`. * *Meaning:* AI or heuristic generation, not manually verified. * *Action:* Flag these in the JSON with `status: "guessed"` for lower confidence scores. 3. **Compound Nouns**: * Code: `compoundN (mkN "costume") "zoot"` * *Challenge:* The lemma is split across multiple strings. * *Action:* Concatenate string literals or take the head noun (first argument) depending on harvester sophistication. ================================================================================================ FILE: docs/Technical-Reference/GF_ARCHITECTURE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 87c07e5535680442a5c7c0985b0f210d62209bd30ca9a02c23d365403638b897 CONTENT_BYTES: 9027 ================================================================================================ # GF Architecture & Developer Guide **Version:** 2.3 **Last Updated:** 2026-02-20 **Context:** SemantiK Architect (Python + Grammatical Framework) ## 1. Architectural Overview This system bridges two fundamentally different paradigms: 1. **Python Domain Layer:** Dynamic, object-oriented, string-heavy. 2. **GF Engine Layer:** Statically typed, functional abstract syntax trees. The critical challenge is the **type mismatch**. Python sees `"Marie Curie"` as a string. GF sees a typed term (e.g., `NP`, `PN`, etc.). Our architecture uses a **Bridge Pattern**: raw runtime data is wrapped into safe GF types before linearization. ### The Stack | Layer | File/Component | Responsibility | | --- | --- | --- | | **Abstract** | `gf/semantik_architect.gf` | Defines the API contract (schema). | | **Concrete (App Grammars)** | `gf/Wiki{WikiCode}.gf` | Implements the schema using the RGL (one per language). | | **Bridge (Syntax Instances)** | `generated/src/Syntax{RglCode}.gf` | Provides `Syntax{RglCode}` instances used by app grammars. | | **Library** | GF Resource Grammar Library (RGL) | Linguistic primitives (`mkS`, `mkCl`, etc.). | | **Adapter** | `app/adapters/engines/gf_wrapper.py` | Converts Pydantic objects → GF trees (PGF expressions). | --- ## 2. Naming: ISO vs WikiCode vs RglCode (Source of Most Bugs) ### Canonical rule: naming is driven by `data/config/iso_to_wiki.json` Confirmed mappings in this repo: - `en` → `{"wiki": "Eng"}` → app grammar is `gf/WikiEng.gf` - `fr` → `{"wiki": "Fre"}` → app grammar is `gf/WikiFre.gf` **Implication:** this codebase does **not** use `WikiEn.gf` / `WikiFr.gf`. ISO-2 codes do not appear directly in grammar filenames. ### Terms - **ISO**: `en`, `fr`, `de` (ISO-639-1) - **WikiCode**: `Eng`, `Fre`, `Ger`, … (from `iso_to_wiki.json`) - **RglCode**: `Eng`, `Fre`, `Ger`, … (typically matches WikiCode here) --- ## 3. Directory Structure & File Hygiene The system historically used two “generated” roots: - `generated/src` (**canonical**) - `gf/generated/src` (**legacy**) On Windows-mounted filesystems (`/mnt/c/...`), symlinks can be unreliable. The commander supports both, but you must treat **`generated/src` as canonical**. ### ✅ Allowed / Canonical - `gf/semantik_architect.gf` — abstract syntax - `gf/Wiki{WikiCode}.gf` — app concrete grammars (e.g., `gf/WikiEng.gf`, `gf/WikiFre.gf`) - `generated/src/Syntax{RglCode}.gf` — bridge “Syntax instances” (e.g., `generated/src/SyntaxEng.gf`) - `gf/semantik_architect.pgf` — compiled binary (usually git-ignored) - `data/config/iso_to_wiki.json` — authoritative mapping (ISO → WikiCode) ### ⚠️ Allowed but Legacy (should be unified) - `gf/generated/src/*` — legacy generated location Prefer to keep it a symlink to `generated/` if your FS supports symlinks; otherwise keep it synchronized and do not hand-edit. ### ❌ Prohibited / Remove or Avoid Creating - `gf/Wiki.gf` — legacy/ambiguous - `gf/WikiEn.gf`, `gf/WikiFr.gf` — wrong naming convention for this repo - `gf/Symbolic*.gf` — **CRITICAL:** local files can conflict with the RGL’s `Symbolic` modules - Any `*.RGL_BROKEN` variants under include paths — can shadow correct modules depending on search path order **Important:** if you see “Generated dirs distinct” warnings, assume there is risk of stale/shadowing modules until unified. --- ## 4. RGL Version Pinning (Alignment Contract) This project relies on a pinned RGL state to avoid “API drift” errors in `Syntax.gf` and other RGL modules. ### Contract - `builder/orchestrator.py` refuses to compile if `gf-rgl` is not pinned to the expected ref. - `python manage.py align --force` is the canonical way to pin RGL and regenerate Tier-1 bridges/app grammars. ### Practical guidance - Treat `gf-rgl/` as a normal git repo clone (submodule configuration is not assumed). - Pin by **ref** (tag/branch/commit), not a “magic commit prefix”. - Prefer a stable ref such as `release-3.12` (or the project’s configured ref), overridable via `SEMANTIK_ARCHITECT_RGL_REF`. --- ## 5. The “Safe” RGL API (Reference) We restrict our RGL usage to a stable subset. ### Core Semantic Constructors | Function | Signature | Description | | --- | --- | --- | | `mkS` | `Cl -> S` | Clause → Sentence | | `mkCl` | `NP -> VP -> Cl` | Predication (“John walks”) | | `mkNP` | `Det -> N -> NP` | Determination (“the animal”) | | `mkVP` | `V2 -> NP -> VP` or `VP -> NP -> VP` | Transitive VP | | `mkAP` | `A -> AP` | Adjectival phrase | ### Structural Helpers | Function | Type | Usage | | --- | --- | --- | | `and_Conj` | `Conj` | List conjunction | | `in_Prep` | `Prep` | “in” | | `symb` | `String -> NP` | **Type bridge** for raw strings | > Note: the exact module providing `symb` depends on RGL version; do not shadow `Symbolic*` locally. --- ## 6. Implementation Rules (Anti-Crash / Stability Rules) These rules exist because some GF compiler + RGL macro patterns can produce unstable compilation behavior. ### Rule #1: Inlining Rule (Scope Safety) Avoid `let` inside `lin` rules when composing complex RGL macros. **Bad:** ```haskell mkEvent subj obj = let v = mkV "participate" in mkS (mkCl subj (mkVP v obj)) ```` **Good:** ```haskell mkEvent subj obj = mkS (mkCl subj (mkVP (mkV "participate") obj)) ``` ### Rule #2: Symbolic Rule (Runtime String Safety) Do not run morphology over runtime variables with `mkPN`/`mkN` when the input is unknown at compile time. **Bad:** `mkLiteral s = mkNP (mkPN s)` **Good:** `mkLiteral s = symb s` ### Rule #3: Modifier Type Rule If you define a `Modifier`, use `Adv` (not `AdV`) unless you have a controlled, language-specific reason. --- ## 7. The Python Adapter Pattern The Python wrapper (`gf_wrapper.py`) must construct ASTs using grammar bridge functions, not raw strings directly. ```python # Wrong: raw strings passed directly # pgf.Expr("mkBio", ["Marie", "Physicist"]) # Right: wrap strings using bridge constructors that exist in semantik_architect.gf subj = pgf.Expr("mkLiteral", [pgf.readExpr('"Marie"')]) prop = pgf.Expr("mkStrProperty", [pgf.readExpr('"Physicist"')]) expr = pgf.Expr("mkBio", [subj, prop]) ``` **Rule:** the adapter must match the exact function names + arities in `gf/semantik_architect.gf`. --- ## 8. Troubleshooting Dictionary | Symptom / Error | Diagnosis | Fix | | -------------------------------------------------------------------------- | --------------------------------------------------- | --------------------------------------------------------------------- | | `Cannot connect to the Docker daemon at unix:///var/run/docker.sock` (WSL) | Linux docker socket not connected to Docker Desktop | Use `docker.exe` from WSL or enable Docker Desktop WSL integration. | | `gf-rgl is not pinned to the expected ...` | RGL pin mismatch | `python manage.py align --force` then build. | | `SyntaxX.gf does not exist` | Missing bridge instance for a Tier language | `python tools/bootstrap_tier1.py --force` (or `manage.py align`). | | `atomic term conflict` / `Symbolic*` conflicts | Local `Symbolic*.gf` shadowing RGL | Delete local `gf/Symbolic*.gf`. | | `Function ... not found` | PGF stale or wrong grammar set compiled | Rebuild PGF and restart API. | | “Generated dirs distinct” warnings | `generated/src` and `gf/generated/src` diverged | Prefer `generated/src`; unify/sync legacy; avoid editing legacy path. | --- ## 9. Build Commands (Current Canonical Flow) ### Always build in WSL (recommended) The default `manage.py` configuration uses Linux venv paths (`venv/bin/python`). Running `manage.py build` in Windows PowerShell will fail unless you provide a Windows venv layout and/or a Windows-aware venv resolver. ### Minimal “Known Good” Flow (Tier-1, small set) ```bash # 1) Align (pins gf-rgl + generates Tier-1 bridges/app grammars) python manage.py align --force # 2) Build only a small language set while iterating python manage.py build --langs en fr ``` ### If you need to refresh RGL inventory for the matrix ```bash python tools/everything_matrix/build_index.py --regen-rgl python tools/bootstrap_tier1.py --force ``` ### Cleaning Do **not** delete all `gf/Wiki*.gf` anymore—those are canonical app grammars in this repo (WikiCode naming). Prefer: ```bash python manage.py clean ``` If you must do a manual clean, focus on compiled artifacts (`*.gfo`, `*.pgf`) and generated outputs, not source grammars. ================================================================================================ FILE: docs/Technical-Reference/GF_Concepts.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 5c3f39f7a6d9c75076160bbc559836509593d34a289f31a34e8f155734395e69 CONTENT_BYTES: 4206 ================================================================================================ # Grammatical Framework (GF): Core Concepts & Compilation ## 1. The Compilation Process (`gf -make`) When we run the command `gf -make`, we are transforming human-readable source code into a machine-optimized format. This process acts like a funnel, merging logic, rules, and vocabulary into a single executable brain. ### The Source (Inputs) We start with three distinct types of `.gf` source files: * **The Logic (`semantik_architect.gf`):** The "Blueprint". It defines *what* can be said (semantics), such as `mkBio` taking an Entity and a Profession. It contains no language-specific words, only mathematical structure. * **The Rules (`WikiEng.gf`):** The "Translator". It implements the blueprint for a specific language (English), defining word order, agreement, and syntax. * **The Vocabulary (`WikiLexiconEng.gf`):** The "Dictionary". It provides the raw words (strings) that plug into the grammar. ### The Action (Compilation) The compiler acts as a processor. It parses all source files, validates them for logical consistency (type checking), and merges them. It ensures that every function defined in the Abstract grammar is correctly implemented in the Concrete grammar. ### The Destination (Output) The result is a single binary file: **`Wiki.pgf`**. * **Binary:** It is optimized for machines, not humans. * **Fast:** It is designed to be loaded instantly into memory (RAM). --- ## 2. The Portable Grammar Format (PGF) **PGF** stands for **Portable Grammar Format**. It is the standard executable format for the GF ecosystem, analogous to Java Bytecode or a compiled binary. ### Why PGF? 1. **Portable:** A single `.pgf` file can be run on any platform—Python, C, Android, or JavaScript runtimes. 2. **Multilingual:** A single `.pgf` file contains *all* compiled languages (English, French, Russian, etc.). This allows the runtime to switch between languages instantly or translate between them without losing meaning. 3. **Efficiency:** Because the grammar is pre-compiled into a mathematical graph, the runtime engine does not need to parse text files. It performs generation and parsing operations in milliseconds. --- ## 3. The "Pivot" Concept A common misconception is that English serves as the central translation language. In GF, **English is not the pivot.** ### The Abstract Syntax Tree (AST) The true pivot is the **Abstract Grammar**. * **Abstract:** $mkBio(Q42, Physicist)$ * **Concrete (English):** "Douglas Adams is a physicist" * **Concrete (French):** "Douglas Adams est un physicien" The system does not translate *English $\rightarrow$ French*. Instead, it parses English into the language-neutral **Abstract** structure, and then linearizes that Abstract structure into French. This ensures that the underlying semantic meaning remains pure, regardless of the surface language. ``` ### Pour l'ajouter rapidement via votre terminal : Vous pouvez copier-coller cette commande dans votre terminal WSL pour créer le fichier d'un coup : ```bash cat < doc/GF_Concepts.md # Grammatical Framework (GF): Core Concepts & Compilation ## 1. The Compilation Process ('gf -make') When we run the command 'gf -make', we are transforming human-readable source code into a machine-optimized format. ### The Source (Inputs) * **The Logic (Abstract):** The Blueprint. Defines *what* can be said. * **The Rules (Concrete):** The Translator. Defines grammar rules (English). * **The Vocabulary (Lexicon):** The Dictionary. Raw words. ### The Action (Compilation) The compiler merges these files, validating logical consistency and type safety. ### The Destination (Output) The result is a single binary file: **Wiki.pgf**. It is machine-optimized for instant loading into RAM. --- ## 2. The Portable Grammar Format (PGF) PGF is the executable format for GF. * **Portable:** Runs on Python, C, JS, Android. * **Multilingual:** Contains ALL languages in one file. * **Efficient:** Binary format allows millisecond generation. --- ## 3. The "Pivot" Concept **English is not the pivot.** The pivot is the **Abstract Grammar**. The system translates via the language-neutral Abstract Syntax Tree (AST), not by translating English to French directly. EOF ``` ================================================================================================ FILE: docs/Technical-Reference/REFERENCE_LINGUISTICS.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: cf275afe945a1361d45ce90f554a1aec39ae88f23c07023fe02f86a04be7757b CONTENT_BYTES: 22112 ================================================================================================ # THEORY_NOTES.md SemantiK Architect – Theoretical Positioning This document explains how the architecture in this repository relates to existing ideas in NLG and linguistic theory. It is **not** an implementation spec, API contract, or schema definition. It is a conceptual map of: * where the design is coming from, * what theoretical traditions it is compatible with, * what kind of multilingual NLG system it is trying to become. It also clarifies a central architectural point: > SemantiK Architect is **not** fundamentally a biography generator. Biography is only one early domain. The intended architecture is **planner-first**, **construction-centered**, **lexically mediated**, and **multi-backend**. The core runtime picture is: ```text semantic input -> normalization -> frame-to-construction bridging -> planning -> ConstructionPlan -> lexical resolution -> realization backend -> SurfaceResult -> API response ``` That theoretical stance matters because it defines what the system is allowed to become as it scales across domains, languages, and realization technologies. --- ## 1. Purpose SemantiK Architect is designed as a **practical multilingual NLG stack** for Abstract Wikipedia and related structured-content workflows. Its internal design is informed by: * **grammar engineering** * **construction grammar** * **frame semantics** * **abstract semantic representations** * **typological and language-family modeling** * **hybrid symbolic generation architectures** * **microplanning / discourse-aware NLG** The goal is to be: * **engineered enough** for large-scale deployment, * **theory-aware enough** to stay compatible with research-grade ideas, * **modular enough** to support many languages and multiple realization backends. The project is therefore best understood as a **construction-centered runtime** with: * semantic inputs, * planner/discourse decisions, * reusable constructions, * explicit slot mapping, * lexical resolution, * family-aware realization, * optional GF-based realization where it is strong. --- ## 2. High-level analogy: where this sits Very roughly, SemantiK Architect sits near several known traditions, without being reducible to any one of them. * Like **Grammatical Framework (GF)**: * it separates language-independent structural intent from language-specific realization, * it treats many sentence patterns as reusable across languages, * it benefits from explicit abstract-to-concrete separation. * Like a **Grammar Matrix** style project: * it factors cross-language regularities into reusable machinery, * it assumes many languages share deeper structural behavior, * it uses configuration and structured linguistic data rather than rewriting everything per language. * Like **construction grammar** and **frame semantics**: * it treats recurrent clause and sentence patterns as reusable constructions, * it assumes meaning is realized through structural packaging, not isolated words alone, * it allows the same content to surface differently across constructions and languages. * Like **AMR / UMR / Ninai-style abstract representations**: * it assumes there is a meaning layer distinct from wording, * it allows input notation to evolve independently from realization code, * it treats normalization and bridging as explicit architectural responsibilities. What SemantiK Architect does **not** try to do is reproduce any one of these traditions in pure form. Instead, it borrows their core strengths: * separation of concerns, * explicit interfaces, * reusable abstractions, * cross-linguistic scaling, * traceable runtime structure. --- ## 3. What kind of system this is At the theoretical level, the intended architecture is: ```text semantic input -> normalization -> frame-to-construction bridge -> planning -> ConstructionPlan / slot_map -> lexical resolution -> realization backend -> SurfaceResult ``` This is important. The central unit is **not** a bio payload and **not** a specific grammar engine. The central unit is a **planned constructional sentence**: a sentence whose semantic roles, information structure, and construction choice have already been determined before realization begins. That means the system should be understood as: * **planner-first**, not renderer-first, * **construction-centric**, not domain-centric, * **backend-agnostic**, not GF-only, * **family-scalable**, not per-language handcrafted only, * **lexically mediated**, not raw-string driven. --- ## 4. Relation to specific ideas and traditions ### 4.1 Grammatical Framework (GF) GF separates: * **abstract syntax** * language-independent structures, * typed constructors, * compositional meaning-to-form mapping, from: * **concrete syntax** * language-specific realization, * morphology, * word order, * agreement. SemantiK Architect is strongly compatible with that way of thinking. In SKA terms, the nearest equivalents are: * normalized semantic frames, * construction classes, * planner output, * `ConstructionPlan`, * language-family and language-specific realization logic. #### Similarities to GF * There is an effort to keep meaning separate from realization. * There is a desire to reuse structural patterns across languages. * There is room for language-specific grammars or realizers. * GF can function as one realization backend. #### Differences from GF * SKA is not built around a single abstract-syntax formalism. * SKA allows multiple realization backends, not one privileged formalism. * SKA uses Python, structured runtime objects, and explicit contracts as primary engineering media. * SKA is intended to stay accessible to mixed teams of engineers and linguistically informed contributors. So the correct theoretical stance is: > SKA is **GF-compatible in spirit**, but not a pure GF system. GF is best treated as: * a powerful grammar-engineering tool, * a strong realization backend for some constructions and languages, * not the sole architectural center of the system. --- ### 4.2 Grammar Matrix and configurable grammar engineering Broad-coverage grammar matrix projects usually separate: * a cross-linguistic core, * a structured parameter space, * language-specific configurations and lexica. This is one of the strongest analogies for SKA. In SemantiK Architect: * family engines act like reusable realization sketches, * language cards and configs act like parameter sets, * constructions act like reusable structural templates, * lexical subsystems capture per-language and per-lexeme variation. SKA is not a full HPSG-style grammar matrix and not a typed feature-structure workbench. But it is aligned with the grammar-matrix idea in one important sense: > many languages should be derivable from shared machinery plus structured variation. That is one of the project’s central scaling ideas. --- ### 4.3 Construction Grammar The closest linguistic affinity of the architecture is probably **construction grammar**. Why: * sentence patterns are treated as reusable units, * constructions package structure and discourse choices together, * the same semantic content can be realized via different constructions, * reusable sentence logic lives in a construction inventory rather than in flat templates or individual lexemes. Examples of constructional thinking in SKA include: * equative and classificatory patterns, * attributive copular patterns, * locatives, * existentials, * possession structures, * relative clauses, * topic-comment structures, * eventive clause patterns, * biography-lead patterns as one construction family among many. This matters theoretically because it means the system is not best described as: * a lexicon plus morphology stack, * a flat template engine, * or raw semantic frames mapped directly to wording. A better description is: > **frame-informed constructional NLG** Frames provide content and semantic roles. Constructions decide how that content is packaged as a sentence. --- ### 4.4 Frame Semantics Frame semantics is relevant because SKA assumes that structured semantic roles matter. The architecture expects inputs that distinguish things like: * actor / patient, * possessor / possessed, * theme / location, * subject / predicate nominal, * topic / focus, * event / participant / circumstance. That is close to the core intuition of frame semantics: * meaning comes with participant structure, * grammatical realization depends on that structure, * multiple surface forms may express overlapping frame content. However, SKA uses frame semantics in a practical engineering sense: * roles are simplified, * frame objects are designed for runtime use, * internal structures need not mirror any published framebank exactly. So the correct claim is: > SKA is **compatible with frame semantics**, but not bound to a single external frame inventory. --- ### 4.5 AMR, UMR, Ninai, and other abstract notations Abstract semantic notations matter because the architecture assumes: * meaning can be represented before wording, * input notation can evolve, * realization should not be permanently tied to one external formalism. That makes systems such as: * Ninai, * UMR, * AMR-like graphs, * Abstract Wikipedia internal semantic forms, relevant upstream inputs. Their correct place in SKA is **before planning**. That means: * external semantic notations should be normalized, * planning should operate on normalized semantic content, * realization should consume a sentence-level constructional plan rather than raw external notation. The theoretical position is: > input formalisms are replaceable; > the constructional runtime architecture should remain stable. That is a strong commitment to separation of concerns. --- ## 5. Internal abstractions and why they look like this ### 5.1 Family engines A major design assumption is that many languages share **deep structural tendencies**. Not perfectly, and not without exceptions, but enough to justify reusable family-level logic. Examples include: * analytic vs fusional vs agglutinative tendencies, * case-heavy vs adposition-heavy marking, * noun-class agreement, * topic-prominent packaging, * article systems, * adjective placement patterns, * possession strategies, * relative clause strategies. These families are partly genealogical, partly typological, and above all **engineering abstractions**. This is not a claim that every language in a family behaves the same. It is a claim that: > many realization decisions can be shared above the individual-language level. That is essential for scale. So family engines are theoretically justified as: * a practical typological abstraction layer, * a middle ground between universalism and per-language handcrafting, * a reusable realization layer behind one construction runtime contract. --- ### 5.2 Constructions vs engines This split is fundamental. * **Constructions** decide what sentence configuration is needed. * **Engines / realizers** decide how that configuration is expressed in a given language. Constructions answer questions like: * Is this an equative? * Is this a classification? * Is this an attributive copular clause? * Is this a locative? * Is this an existential? * Is this a possession structure? * Is this a topic-comment clause? * Is this a bio-lead identity sentence? Engines answer questions like: * Is there an article? * How is agreement marked? * What is the default word order? * How is possession expressed in this language? * How are topic and focus surfaced? * What morphology or function words are required? Without this split, multilingual generation collapses into either: * too much language-specific logic inside every construction, * or too much semantic logic hidden inside realization backends. So this separation is both: * linguistically motivated, * architecturally necessary. --- ### 5.3 Planning and discourse The architecture also implies a planner/discourse layer. This matters because sentence generation is not just: * selecting words, * inflecting them, * placing them in order. It also includes: * deciding which construction to use, * choosing canonical vs topic-prominent packaging, * deciding which entity is discourse-prominent, * deciding which role should be foregrounded, * determining how the same semantic content should become a sentence. This makes the system compatible with ideas from: * information structure, * discourse planning, * centering-style approaches, * microplanning in NLG. The system does not need a full discourse theory to justify this. Even a light planner already changes the architecture substantially. The key theoretical point is: > planning belongs between semantics and realization. --- ### 5.4 Slot mapping as an explicit layer The architecture also benefits from making **slot mapping** explicit. A construction is not just a label. It expects semantically named inputs such as: * `subject` * `predicate_nominal` * `predicate_adjective` * `location` * `agent` * `patient` * `theme` * `topic` * `comment` This matters theoretically because it avoids collapsing: * frame roles, * construction roles, * lexical items, * backend arguments into one undifferentiated structure. Explicit slot mapping makes it clearer that: * semantics provides role content, * constructions define the packaging, * realization consumes already-packaged inputs. That is good both linguistically and architecturally. --- ### 5.5 Lexical resolution as a separate layer The architecture also assumes that lexical choice and lexical normalization must not be hidden inside whichever renderer happens to run. This is theoretically important because: * semantic content is not identical to lexical form, * raw strings are not enough for high-quality multilingual realization, * lexical provenance matters, * lexical uncertainty matters, * language-specific realization often depends on more than a label. So lexical resolution is not just preprocessing. It is a proper layer between construction planning and surface realization. That makes the system more compatible with both: * classical lexicalist insights, * practical multilingual NLG requirements. --- ### 5.6 Realization backends as interchangeable surface technologies Another important abstraction is that realization technology is not the same thing as architecture. SKA allows multiple backends, such as: * family-oriented realizers, * GF-based realizers, * safe-mode fallback realization, * future hybrid systems. This means the system’s conceptual center cannot be any one backend. The stable architectural object is the construction-level plan, not the backend’s internal representation. So the correct theoretical reading is: > realization backends are interchangeable surface technologies operating over one shared construction runtime. --- ## 6. What this system is not It is useful to state clearly what SemantiK Architect is **not** trying to be. ### 6.1 Not a pure template system It can contain templates and reusable surface patterns, but the intended architecture is richer than static string templating. ### 6.2 Not a pure semantic formalism It is not trying to be a full logical language or graph formalism in itself. ### 6.3 Not a pure grammar formalism It is not a GF clone, not an HPSG workbench, not an LFG implementation, and not a single-formalism grammar laboratory. ### 6.4 Not an LLM-only generation stack Learned components may become useful later, but the architecture is fundamentally built around explicit structure, traceability, and deterministic runtime behavior. ### 6.5 Not a biography-only system Biography is one early domain and one useful test case. It must not become the hidden architectural center. This point is crucial. ### 6.6 Not a renderer-first architecture No renderer should become the place where sentence meaning, construction choice, and discourse packaging are secretly decided. That would undo the central architectural separation. --- ## 7. Core theoretical tradeoffs ### 7.1 Expressiveness vs maintainability The architecture deliberately avoids: * maximal formal elegance, * maximal notational purity, * maximal linguistic detail everywhere. Instead it chooses: * explicit layers, * typed-enough structures, * configurable data, * readable runtime code, * testable constructions, * debuggable contracts. This is a standard engineering tradeoff: less theoretical purity, more maintainable multilingual infrastructure. --- ### 7.2 Family-level generalization vs per-language accuracy Family engines risk overgeneralization. That risk is real. But the alternative is also costly: fully bespoke logic per language does not scale. So the system adopts a pragmatic compromise: * share what can be shared, * isolate what must be language-specific, * allow override points, * keep capability tiering explicit. The theory here is practical: typological reuse is worth it if the override model is real. --- ### 7.3 Planner-first vs renderer-first architecture This is one of the most important theoretical choices in the repository. A renderer-first system tends to collapse: * semantics, * construction choice, * lexical assumptions, * and wording into one backend. A planner-first system keeps them apart. SKA is more coherent when understood as planner-first. That means: * sentence structure should be selected before realization, * renderers should realize plans, not invent them, * no backend should become the semantic center of the system. --- ### 7.4 Generic runtime vs domain-specific shortcuts There is always pressure in multilingual systems to special-case a successful early domain. Biography is the obvious example. Theoretical caution: * domain-specific shortcuts are tempting, * but when they become architecture, they distort the whole system. So early domains should be treated as: * motivating examples, * coverage targets, * not architectural masters. --- ### 7.5 Strong contracts vs implementation flexibility The architecture benefits from strong shared contracts: * normalized frames, * `ConstructionPlan`, * `slot_map`, * lexical references, * `SurfaceResult`. At the same time, it needs implementation flexibility underneath those contracts. This tradeoff is central: * contracts should be stable enough to prevent architectural drift, * implementations should be flexible enough to support multiple backends and gradual migration. That combination is what allows both rigor and evolution. --- ## 8. Theoretical view of the target runtime The most coherent theoretical reading of the intended system is: ### 8.1 Semantic layer Represents content and role structure. ### 8.2 Normalization layer Converts upstream payloads or notations into stable internal frame objects. ### 8.3 Frame-to-construction bridge Maps normalized semantic content toward construction-oriented planning. ### 8.4 Planning layer Selects information packaging, topic/focus behavior, and construction choice. ### 8.5 Construction layer Defines reusable sentence patterns independent of any single language. ### 8.6 Slot-mapping layer Assigns construction-specific semantic inputs into explicit realization slots. ### 8.7 Lexical resolution layer Connects planned slots to lexical identities, lexical features, and provenance. ### 8.8 Realization layer Implements family- and language-specific wording, morphology, and surface order. ### 8.9 Backend layer Allows different realization technologies: * GF, * family engines, * safe-mode fallback, * future hybrid systems. This layered reading best matches both: * the design intent, * the practical needs of a multilingual NLG system, * the need to keep architecture stable while implementations evolve. --- ## 9. Future theoretical directions This architecture can grow in several research-compatible directions. ### 9.1 Richer frame inventories Expand beyond narrow early domains into broader semantic frame families. ### 9.2 Stronger construction inventories Make construction classes, slot contracts, and construction capabilities more explicit and reusable. ### 9.3 Better discourse models Add more principled topic, salience, anaphora, aggregation, and sentence-ordering models. ### 9.4 Stronger lexical semantics Use richer lexical references, lexical typing, provenance tracking, and confidence-aware fallback. ### 9.5 Hybrid symbolic and learned systems Use learned components for ranking, lexical choice, or variation, while preserving explicit constructional planning and debuggable runtime structure. ### 9.6 Closer GF and grammar-engineering interoperability Use GF where it is a strong backend, without turning the entire architecture into a single-formalism system. ### 9.7 Stronger contract-centered evaluation Evaluate systems not only by output quality, but also by whether they preserve: * construction identity, * slot integrity, * lexical traceability, * fallback transparency, * backend-independent semantics. --- ## 10. Summary SemantiK Architect should be understood as: * a **construction-centered multilingual NLG architecture**, * informed by **GF**, **grammar matrix thinking**, **construction grammar**, **frame semantics**, and **abstract semantic formalisms**, * implemented in a pragmatic, runtime-oriented way, * with explicit separation between: * semantics, * normalization, * planning, * constructions, * slot mapping, * lexical resolution, * realization. Its theoretical identity is therefore not: * just templates, * not just GF, * not just semantic frames, * not just language-family rules, but rather: > a practical multilingual NLG system whose central abstraction is the > **planned constructional sentence**, expressed as a stable construction-level plan > and realized through configurable, family-aware, backend-agnostic mechanisms. These notes exist to make that theoretical identity explicit, so the system can evolve without losing its conceptual center. ================================================================================================ FILE: docs/Technical-Reference/RESEARCH_GF_MAPPING.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 1ab4ea3c2f0b00f560ceba0708bcabe92731851f9ea1b8028e911fccd90e4e4b CONTENT_BYTES: 20383 ================================================================================================ # GF Integration Design Status: proposed Owner: SKA runtime maintainers Scope: integrate selected parts of **Grammatical Framework (GF)** into **SemantiK Architect (SKA)** while preserving SKA’s planner-first, construction-centered runtime. --- ## 1. Purpose This document defines how GF fits into SKA’s runtime and tooling model. The key decision is: > **GF is a renderer backend and an offline enrichment source, not the architecture itself.** This document exists to remove drift between: * the documented planner/construction architecture, * the current runtime generation stack, * GF-backed realization, * family-engine realization, * safe-mode fallback, * offline GF-derived morphology, syntax, and QA assets. It makes explicit what is already true in the repository: * GF/PGF is already part of the generation stack, * family engines remain first-class, * lexical resolution remains a shared SKA concern, * runtime authority belongs to planner + construction contracts, not to a backend. --- ## 2. Goals and Non-Goals ## 2.1 Goals 1. **Use GF where it actually fits** * as an **optional runtime renderer backend** for selected constructions and languages, * as an **offline source** of morphology paradigms, syntactic patterns, and QA examples, * as a **grammar-engineering reference** for improving SKA’s constructions and family engines. 2. **Keep SKA’s architecture intact** SKA remains centered on: * semantic frames, * discourse/planning, * construction selection, * construction plans, * lexical resolution, * family engines, * language/morphology data. 3. **Keep planning authoritative** Runtime generation remains: ```text frame -> planner -> construction plan -> lexical resolution -> renderer backend -> surface result ``` GF consumes the same runtime contract as other renderers. 4. **Support mixed backend capability** * some languages/constructions may use GF, * some may use family engines, * some may fall back to safe mode. All three must coexist without semantic drift. 5. **Make integration auditable** * clear capability metadata, * clear runtime traces when GF is selected, * versioned and reviewable offline imports, * explicit provenance on GF-derived artifacts. 6. **Keep GF optional** * runtime does not require GF unless the selected backend is GF, * offline tooling is isolated, * non-GF languages continue to work through family or safe-mode backends. --- ## 2.2 Non-Goals This design does **not**: * make GF the core architecture of SKA, * make all runtime generation dependent on GF, * adopt GF ASTs as SKA’s primary internal contract, * require every language to have a GF runtime path, * replace family engines, lexicon resolution, or JSON language/morphology data. This design also does **not** require: * parsing, * reversible grammar support, * full concrete GF coverage for every language, * direct exposure of GF ASTs to API clients. --- ## 3. Executive Summary GF has two valid roles in SKA: 1. **Runtime role** * GF is a renderer backend for selected `(lang_code, construction_id)` pairs. 2. **Offline role** * GF resources can be harvested into morphology data, syntax notes, QA material, and coverage/provenance artifacts. GF is therefore: * **not** the architecture, * **not** offline-only, * **not** mandatory, * **not** the semantic or planning layer. The authoritative SKA runtime remains: ```text API payload -> frame normalization -> planner -> construction plan -> lexical resolution -> renderer backend -> surface result ``` GF participates inside the **renderer backend** layer at runtime and inside the **offline enrichment** layer outside runtime. --- ## 4. Current Architectural Position SKA already contains the main ingredients required for this integration: * semantic/domain frames, * discourse/planning, * `PlannedSentence`, * `construction_id`, * construction inventory, * family-oriented engines, * lexicon and lexical normalization, * GF integration in the current generation stack. The current architectural tension is that the repository contains both: * the intended planner/construction architecture, and * a still-existing direct generation path that can bypass the shared runtime contract. This document resolves that tension by making GF subordinate to the same runtime boundary used by other backends. --- ## 5. Architectural Decision ## 5.1 Authoritative runtime contract The authoritative runtime flow is: ```text API payload -> normalized frame -> planner -> construction plan -> lexical resolution -> renderer backend -> surface result ``` This implies: * GF must consume a `ConstructionPlan`, not raw API payloads as its primary contract. * family engines must consume the same plan. * safe mode must consume the same plan. * direct frame-to-GF generation is compatibility behavior only. ## 5.2 GF’s place in the runtime GF sits in the **renderer layer**. GF is **not**: * the planner, * the semantic model, * the API contract, * the lexicon subsystem, * the source of construction truth. GF **is**: * a renderer backend for supported `(lang_code, construction_id)` pairs. ## 5.3 GF’s place in offline tooling GF also sits in the **offline enrichment layer**. GF can provide: * morphology paradigms, * syntax references, * example sets, * test material, * coverage data, * provenance metadata. --- ## 6. Canonical Runtime Boundary for GF GF integration must use the same canonical runtime objects as the rest of Batch 1. ## 6.1 Renderer input GF consumes a `ConstructionPlan` with at least: * `construction_id` * `lang_code` * `slot_map` * optional `topic_entity_id` * optional `focus_role` * `metadata` Renderer-safe realization options belong under: ```python metadata["generation_options"] ``` Typical option families include: * `tense` * `aspect` * `polarity` * `register` * `definiteness` * `voice` * `style` * `allow_fallback` * `force_backend` * `debug` GF may also consume normalized lexical content already prepared by lexical resolution. ## 6.2 Renderer output GF returns a `SurfaceResult` with: * `text` * `lang_code` * `construction_id` * `renderer_backend = "gf"` * `debug_info` Required debug keys include: * `construction_id` * `renderer_backend` * `lang_code` * `fallback_used` Typical optional GF-specific debug fields include: * `resolved_language` * `gf_function` * `ast` * `backend_trace` * `warnings` * `timings_ms` ## 6.3 Boundary rule GF may choose: * backend-specific AST construction, * concrete grammar selection, * local linearization strategy. GF may **not**: * change `construction_id`, * reinterpret slot meanings, * invent missing semantic structure, * become the planner by stealth. --- ## 7. GF Integration Modes GF is integrated in two distinct modes. ## 7.1 Mode A — Runtime renderer backend In this mode, GF is used during request-time generation. Requirements: * compiled PGF is available, * `(lang_code, construction_id)` capability is declared, * renderer dispatch selects GF, * the GF adapter consumes the shared `ConstructionPlan` contract. Typical use: * high-fidelity realization for selected constructions, * controlled multilingual generation, * deterministic output where GF support is strong. ## 7.2 Mode B — Offline knowledge provider In this mode, GF is not used in request-time generation. Requirements: * GF tools are available offline or in CI, * export/harvest scripts produce intermediate artifacts, * SKA conversion tools transform them into JSON/CSV/Markdown artifacts, * provenance is preserved. Typical use: * morphology enrichment, * syntax notes, * construction review, * test generation, * regression support. --- ## 8. High-Level Integration Architecture ## 8.1 Planning layer Authoritative planner-side responsibilities: * normalized frames, * discourse/planning, * construction selection, * `construction_id`, * sentence packaging, * topic/focus metadata. ## 8.2 Construction runtime contract Shared runtime objects include: * `PlannedSentence` * `ConstructionPlan` * `SlotMap` * `EntityRef` * `LexemeRef` * `SurfaceResult` This layer is the required boundary between planning and realization. ## 8.3 Renderer layer Runtime renderer backends are: * family backend, * GF backend, * safe-mode backend. All must implement the same logical runtime surface: ```python async def realize(construction_plan: ConstructionPlan) -> SurfaceResult ``` ## 8.4 Offline GF layer Offline GF tooling includes: * export scripts, * conversion scripts, * provenance capture, * import into JSON/CSV/docs, * capability review material, * QA artifact generation. --- ## 9. Runtime Data Flow ## 9.1 Runtime path ```text Request JSON -> semantic frame normalization -> planner -> construction plan -> lexical resolution -> renderer dispatch -> gf backend -> family backend -> safe_mode backend -> surface result -> API response ``` ## 9.2 Offline path ```text GF grammars / RGL / examples -> export scripts -> intermediate GF artifacts -> SKA conversion tools -> morphology data / syntax notes / QA datasets / provenance docs ``` ## 9.3 Isolation rule * runtime does not require GF unless GF is selected, * languages without GF support continue to work, * offline GF tooling remains isolated from normal runtime execution, * generated artifacts are stored as normal SKA assets, not as hidden backend state. --- ## 10. Backend Selection Policy ## 10.1 Selection inputs Renderer dispatch should be based on: * `(lang_code, construction_id)` capability, * backend health/readiness, * configuration flags, * allowed fallback policy, * explicit debug/test overrides when requested. ## 10.2 Recommended priority 1. **GF backend** * when the requested `(lang_code, construction_id)` is supported and healthy. 2. **Family backend** * when family realization is supported. 3. **Safe-mode backend** * when deterministic fallback is allowed. ## 10.3 Semantic guarantee Backend choice may change: * phrasing quality, * morphology richness, * debug detail, * local idiomatic realization. Backend choice must **not** silently change: * `construction_id`, * semantic role assignment, * truth-conditional content, * planner-authorized information packaging. --- ## 11. GF Runtime Backend Requirements ## 11.1 Input requirements The GF adapter must consume: * `construction_id` * `lang_code` * normalized `slot_map` * lexicalized or lexicalizable slot values * optional planner metadata * `metadata["generation_options"]` The GF adapter must **not** require: * raw API request shape, * router-specific payload hacks, * domain-specific direct frame flattening as its long-term contract. ## 11.2 Output requirements The GF adapter must return a `SurfaceResult` with: * `text` * `lang_code` * `construction_id` * `renderer_backend = "gf"` * truthful `debug_info` Recommended GF debug fields: * `resolved_language` * `gf_function` * `ast` * `slot_keys` * `fallback_used` * `backend_trace` ## 11.3 Legacy compatibility rule If direct frame-to-GF behavior exists in legacy code, it must be treated as: * compatibility-only behavior, * isolated adapter logic, * temporary migration support. It must not remain the canonical runtime boundary. --- ## 12. Construction-Centered GF Capability GF support must be tracked by **construction**, not just by language. Correct capability unit: ```text (lang_code, construction_id) ``` Examples: * `("fr", "copula_equative_simple")` * `("fr", "copula_locative")` * `("en", "topic_comment_eventive")` This matters because: * a language may have GF assets but only partial construction coverage, * different constructions may require different mapping quality, * backend choice must be explicit and testable. GF presence for a language is therefore **not enough** to claim runtime support. --- ## 13. Relationship to Family Engines ## 13.1 Family engines remain first-class GF integration does **not** replace family engines. Family engines remain necessary because they: * scale across larger language inventories, * encode SKA-native morphology/configuration, * work where GF coverage is absent, * support constructions beyond current GF coverage. ## 13.2 Division of responsibility * **Planner / construction layer** * chooses structure and semantic packaging. * **Lexical resolution** * normalizes entities and lexemes for realization. * **GF backend** * realizes supported constructions for supported languages. * **Family backends** * realize supported constructions using family-native morphology and language data. * **Safe-mode backend** * provides deterministic fallback when stronger backends are unavailable. --- ## 14. Relationship to Lexical Resolution GF integration must not collapse lexical handling into a grammar-only model. Lexical resolution remains a shared SKA concern because it is reused by: * multiple constructions, * multiple backends, * multiple languages, * QA/coverage tooling, * entity/lexicon bridges. GF may contribute lexical insights and examples, but runtime lexical normalization remains outside the GF adapter. Practical rule: * renderers consume lexical decisions, * they do not own canonical lexical classification. --- ## 15. Offline Morphology Integration ## 15.1 Objective Use GF morphology resources to enrich SKA morphology data without forcing GF at runtime for every language. ## 15.2 Target outputs GF exports may enrich: * language/family morphology configs, * language cards, * lexicon-side feature data where appropriate, * construction capability metadata. ## 15.3 Intermediate representation Intermediate export files should preserve: * language, * GF version, * module provenance, * category names, * slot inventories, * paradigm examples, * transformation rules. ## 15.4 Conversion responsibilities Conversion tooling should: 1. read GF exports, 2. map GF categories/features to SKA categories/features, 3. generate or merge SKA artifacts, 4. attach provenance metadata, 5. flag unmapped or suspicious items for review. --- ## 16. Offline Syntax Pattern Harvesting ## 16.1 Objective Use GF analyses and examples to improve SKA constructions and family-engine defaults. ## 16.2 What is harvested Potential harvest outputs include: * construction insights, * parameterization patterns, * word-order options, * feature interactions, * curated examples, * regression examples. GF syntax code is **not** imported as SKA’s primary runtime logic. ## 16.3 Deliverables * family-specific syntax notes, * construction adjustments, * updated language/family configuration where justified, * regression example sets. --- ## 17. Offline Test Integration ## 17.1 Objective Use GF grammars and generated examples to create high-value QA material. ## 17.2 Target outputs * JSON/CSV QA rows, * minimal pairs, * agreement tests, * construction-specific regression suites, * backend parity checks where feasible. ## 17.3 Provenance requirements Each GF-derived test artifact should record: * source = GF/RGL * GF version * modules used * language * generation date * transformation script version --- ## 18. Repo Alignment GF integration work should align with the actual runtime adapter and grammar locations in the repository. Relevant runtime/backend paths include: ```text app/adapters/engines/construction_realizer.py app/adapters/engines/family_construction_adapter.py app/adapters/engines/gf_construction_adapter.py app/adapters/engines/safe_mode_construction_adapter.py app/adapters/engines/gf_wrapper.py app/adapters/engines/gf_engine.py app/adapters/engines/python_engine_wrapper.py ``` Relevant contract/runtime docs include: ```text docs/RESEARCH_GF_MAPPING.md docs/contracts/construction_runtime_contract.md docs/contracts/planner_realizer_interfaces.md docs/grammar/construction_renderer_contract.md docs/architecture/construction_runtime_alignment.md docs/architecture/construction_runtime_flow.md ``` Relevant grammar/runtime surface files include: ```text gf/SemantikArchitect.gf gf/WikiI.gf gf/WikiEng.gf gf/WikiFre.gf ``` This document is about how those areas fit together, not about creating a separate parallel architecture. --- ## 19. Implementation Plan ## Milestone 0 — Runtime alignment 1. make the shared construction runtime contract authoritative, 2. ensure planner output is authoritative, 3. adapt GF to consume `ConstructionPlan`, 4. isolate direct frame-to-GF generation as compatibility behavior only. ## Milestone 1 — First canonical runtime GF slice 1. pick one construction family, 2. support one or two languages through the GF backend, 3. add standardized debug/provenance fields, 4. validate semantic parity against family realization. ## Milestone 2 — Offline morphology and syntax harvesting 1. add or formalize GF export/conversion tooling, 2. export one language slice, 3. convert it into SKA-native artifacts, 4. document provenance and review workflow. ## Milestone 3 — QA integration 1. export GF-derived test cases, 2. convert them into SKA QA suites, 3. add construction-aware regression tests, 4. track backend parity where possible. ## Milestone 4 — Broader construction coverage 1. expand GF support by construction, 2. expand capability metadata, 3. expand offline harvest coverage where useful, 4. keep backend behavior explicit and reviewable. --- ## 20. Risks and Mitigations ## 20.1 Risk: GF becomes hidden architecture If more logic drifts into the GF adapter, planner/construction authority erodes. **Mitigation** * keep `ConstructionPlan` authoritative, * keep backend interfaces shared, * forbid backend-private construction semantics. ## 20.2 Risk: mismatch between GF and SKA feature systems GF categories and SKA feature inventories will not always align directly. **Mitigation** * use explicit mapping tables, * allow partial imports, * validate unmapped features, * require human review for promoted changes. ## 20.3 Risk: semantic drift across backends Different backends may realize the same construction differently enough to alter meaning or discourse packaging. **Mitigation** * compare outputs at the construction level, * keep `slot_map` explicit, * require debug traces showing construction and backend. ## 20.4 Risk: maintenance burden GF updates, module changes, or grammar drift may create import or runtime instability. **Mitigation** * record GF version in every derived artifact, * treat offline imports as reviewable diffs, * isolate capability metadata per `(lang_code, construction_id)`. ## 20.5 Risk: overfitting to GF assumptions GF is powerful, but not the only valid linguistic representation for SKA. **Mitigation** * treat GF as one backend and one source, * preserve family-engine and lexicon-centered architecture, * avoid baking GF-specific assumptions into public runtime contracts. --- ## 21. Licensing and Attribution GF-derived artifacts and GF-backed runtime components must preserve attribution metadata. At minimum, track: * source * GF version * relevant modules * license note * generation timestamp Example: ```json { "_meta": { "source": "GF Resource Grammar Library", "license": "BSD-style", "gf_version": "3.12", "generated_at": "2026-03-10T00:00:00Z" } } ``` Runtime debug info may also expose non-sensitive GF provenance when useful for testing or review. --- ## 22. Summary SKA does **not** adopt GF’s architecture as SKA’s architecture. SKA does **use GF in two roles**: 1. **runtime renderer backend** for selected `(lang_code, construction_id)` pairs, 2. **offline knowledge source** for morphology, syntax insights, and QA material. The authoritative SKA runtime remains: ```text frame -> planner -> construction plan -> lexical resolution -> renderer backend -> surface result ``` GF is therefore: * useful, * optional, * construction-aware, * runtime-valid, * architecturally subordinate to SKA’s shared runtime contract. The main design correction relative to the older draft is explicit: * GF is **not** forced into an offline-only role, * GF is already part of the actual generation stack, * but GF still remains **one backend among several**, not the runtime center. ================================================================================================ FILE: docs/Technical-Reference/RGL_DISCOVERY_STRATEGY.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 1078edb280aa1a40efcfef84ebb6bcc61628a4277f95d0736c62be1c8de0862e CONTENT_BYTES: 5434 ================================================================================================ # 🧩 RGL Integration & Dynamic Discovery Strategy **SemantiK Architect — Internal Reference (Everything Matrix / iso2-keyed)** ## 1) The “Naming Mismatch” Problem Integrating the **GF Resource Grammar Library (RGL)** requires bridging three different naming schemes: | Layer | Identifier | Example (French) | Example (German) | Example (Chinese) | |---|---|---|---|---| | **Semantik Architect (System Keys)** | **ISO-639-1 (iso2, lowercase)** | `fr` | `de` | `zh` | | **GF RGL (Module Suffix)** | Legacy 2–3-letter suffix | `Fre` | `Ger` | `Chi` | | **File System** | Folder name | `french` | `german` | `chinese` | **The conflict:** You cannot derive RGL module names by capitalizing iso codes. - `SyntaxFr` (expected) ≠ `SyntaxFre` (actual) - `SyntaxDe` (expected) ≠ `SyntaxGer` (actual) Historically, teams tried hardcoded dictionaries or parsing `languages.csv`, but that creates maintenance burden and `languages.csv` is not a reliable ISO translation layer. --- ## 2) The Solution: “Inventory + Normalization” (Everything Matrix Suite) ### Single orchestrator contract The canonical “refresh” entrypoint is: - `tools/everything_matrix/build_index.py` It produces: - `data/indices/everything_matrix.json` (iso2-keyed) - It **does not** rescan `gf-rgl/src` during a normal run. - It treats `data/indices/rgl_inventory.json` as an **input artifact**. ### RGL discovery is isolated in one scanner (debug tool) The only component that walks `gf-rgl/src` is: - `tools/everything_matrix/rgl_scanner.py` It produces the source-of-truth inventory: - `data/indices/rgl_inventory.json` **Side-effect policy:** - Imported as a library: side-effect free by default - CLI mode can write the inventory when explicitly requested (e.g., `--write`), or when `build_index.py --regen-rgl` is used. ### Canonical normalization A shared normalization module (e.g., `tools/everything_matrix/norm.py`) is the source of truth for mapping: - Wiki suffixes / iso3 aliases → canonical **iso2** - iso2 → display names (when needed) This makes the system deterministic and avoids “Wiki vs ISO” row mismatches. --- ## 3) Resolution Algorithm Used by the Builder When the build system needs to compile a language (example: `fr`): ### Step 0: Normalize the requested language key Normalize any incoming code (`Fre`, `fra`, `fr`) to canonical **iso2** using `config/iso_to_wiki.json`. **Rule:** the system key is always `iso2` lowercase. ### Step 1: Matrix lookup (iso2 → metadata) Read `data/indices/everything_matrix.json` for orchestration metadata and readiness signals. Example (illustrative): ```json "fr": { "meta": { "folder": "french", "origin": "rgl", "tier": 1 } } ```` ### Step 2: Inventory lookup (iso2 → exact RGL module names) Read `data/indices/rgl_inventory.json` for the exact module set and their real on-disk paths. Example (illustrative): ```json "languages": { "fr": { "path": "gf-rgl/src/french", "modules": { "Syntax": "gf-rgl/src/french/SyntaxFre.gf", "Paradigms": "gf-rgl/src/french/ParadigmsFre.gf" }, "blocks": { "CAT": 10, "NOUN": 10, "PARA": 10, "GRAM": 10, "SYN": 10 } } } ``` **Important:** the builder should not glob `Syntax*.gf` during normal operation; it should trust `rgl_inventory.json` for determinism. ### Step 3: Connector generation (the “Empty Connector” pattern) Generate the compatibility bridge using the exact discovered module names from the inventory. Example: ```haskell -- Generated: WikiFre.gf (suffix derived from inventory) concrete WikiFre of SemantikArchitect = WikiI ** open SyntaxFre, ParadigmsFre in { -- Empty body guarantees compilation success. -- Vocabulary is injected at runtime. } ``` **Why empty bodies:** * The old approach tried to emit lexical lines (e.g., `mkNP (mkN "animal")`) and failed for languages with different parameterization. * The new approach guarantees compilation success; runtime lexicon injection handles vocabulary. ### Step 4: Regeneration / self-healing behavior If `rgl_inventory.json` is missing or stale: * Preferred: `python tools/everything_matrix/build_index.py --regen-rgl` * Or debug: `python tools/everything_matrix/rgl_scanner.py --write` --- ## 4) Directory Structure Constraints This approach assumes a stable repo root and paths derived from the canonical config: * Canonical config: `data/config/everything_matrix_config.json` * RGL base path: `gf-rgl/src` (configurable via `rgl_base_path`) Expected layout: ```text /ProjectRoot/ ├── semantik-architect/ │ ├── tools/everything_matrix/build_index.py │ ├── data/config/everything_matrix_config.json │ └── data/indices/rgl_inventory.json └── gf-rgl/ └── src/ ├── french/ │ └── SyntaxFre.gf └── ... ``` --- ## 5) Why this is Robust 1. **Deterministic builds** * The build uses `rgl_inventory.json`, avoiding “moving target” scans during compilation. 2. **Zero hardcoded mapping tables** * The only mapping source is `config/iso_to_wiki.json` and the scanner’s on-disk discovery. 3. **Schema-compatible orchestration** * Everything Matrix keys are canonical iso2, so Zone A/B/C/D land in the same row. 4. **Self-healing via explicit regen** * If RGL changes module suffixes, re-running the scanner regenerates inventory and downstream connectors without code changes. ================================================================================================ FILE: docs/Technical-Reference/TechnicalStatus23dec2025.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: dea99839f0ba5c37ba139fa875bf36b3557e09073aad5e1da7a446fc172addb1 CONTENT_BYTES: 3719 ================================================================================================ ### 📋 Technical Reality Check: SemantiK Architect (v2.0) **Overall Status:** **Functional Beta**. The engine runs, the pipes connects, and the "Brain" (Matrix) sees everything. But the "Content" is spotty, and performance has known bottlenecks. --- ### 🟢 What Actually Works (The Green Zone) **1. The Plumbing (Infrastructure)** * **Hexagonal Isolation:** The separation is real. You can swap the file system or Redis without breaking the core linguistic logic. * **The "Everything Matrix" Scanner:** This is the most mature part. The indexer script successfully scans the file system, and the Next.js frontend successfully visualizes the 14 health blocks. It correctly identifies which languages are broken. * **Tier 1 Grammars (RGL):** For high-resource languages (English, Finnish, French), the system produces industrial-grade text with complex morphology (cases, genders). * **Docker/WSL Hybrid:** The split between Windows (Frontend) and Linux (Backend/GF) is stable and documented. **2. The New "Dual-Path" Fix** * **Prototyping:** As of our last coding session, the API now accepts `UniversalNode`. You can now successfully "hallucinate" new functions (e.g., `mkIsAProperty`) without the validator crashing. This unblocked your drafting phase. --- ### 🔴 What Doesn't Work Yet (The Red Zone) **1. The "Q42" Fallback Problem (Lexicon Gaps)** * **The Issue:** The architecture assumes `morphodict` or `people.json` has every word. It doesn't. * **The Symptom:** When a word is missing, the engine falls back to the raw string. You will see output like *"Q42 lives in Paris"* instead of *"Douglas Adams lives in Paris"*. * **Severity:** High. This makes the output unreadable for end-users in many languages. **2. Cold Start Latency** * **The Issue:** The PGF binary is massive. Loading `semantik_architect.pgf` into RAM takes significant time (10-60s) and consumes 500MB+ per worker. * **The Symptom:** The first request after a deploy hangs. Scaling workers is memory-expensive. **3. Tier 3 Reliability (The AI Gamble)** * **The Issue:** For the ~60 low-resource languages (Tier 3), you rely on the "Architect Agent" (LLM) and "Weighted Topology". * **The Symptom:** AI is probabilistic. It sometimes generates valid GF code, but often requires the "Surgeon" agent to fix it in a loop. It is not "fire and forget"; it is "fire and pray." * **Matrix Reality:** The Matrix shows 79 languages, but the documentation admits only ~12 are "Production Ready". That means **67 languages are likely stubs or broken.** **4. API Security** * **The Issue:** The API is currently wide open. There is no Authorization Bearer layer (OAuth2/JWT) implemented yet. * **The Symptom:** Anyone with access to port 8000 can trigger expensive compilations. --- ### 🟡 The "Kind of Works" (Yellow Zone) **1. Ninai Compatibility** * **Status:** The *Adapter* exists and parses the JSON tree. * **But:** It is currently a strictly mechanical translation. It doesn't yet handle the full semantic nuance of Abstract Wikipedia's Z-Objects, just the structural shape. **2. The "Surgeon" (Self-Healing)** * **Status:** It can catch compilation errors and try to patch them. * **But:** It has a `MAX_RETRIES` limit. If the AI fails 3 times, the language is dropped. It is not a magic wand; it's a retry loop with a smart guesser. ### Summary Verdict You have built a **Ferrari Engine** (GF/Python) inside a **Professional Garage** (Matrix/Docker), but you are currently driving it with **Empty Gas Tanks** (Missing Lexicon) for most languages. **Next Immediate Step:** Focus on **Data Injection** (Lexicon), not more Architecture. You need to fill those empty "Health Blocks" in the Matrix. ================================================================================================ FILE: docs/Technical-Reference/UPGRADE v2-0 Omni-Upgrade Architecture Specification.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: dc4a59ad519aada16e16b8d2ce728d1dd9ad3457765fdcf197357b5541159d01 CONTENT_BYTES: 10216 ================================================================================================ Here is the **Exhaustive v2.0 "Omni-Upgrade" Architecture Specification**. This document is the definitive technical blueprint for the upgrade. It expands on every logic flow, data structure, and integration point derived from the `Ninai` and `Udiron` code audits. You should replace the contents of `docs/13-V2_ARCHITECTURE_SPEC.md` with this version. --- # 🏗️ v2.0 "Omni-Upgrade" Architecture Specification (Exhaustive) **SemantiK Architect** ## 1. Executive Summary This specification defines the "v2.0" architectural expansion. The goal is to evolve the system from a **Sentence-Level Rule-Based Engine** into a **Context-Aware, Interoperable, and AI-Augmented Platform**. **The 7 Core Pillars of Upgrade:** 1. **Interoperability:** Ninai Bridge (JSON Object Adapter). 2. **Standards:** UD Exporter (CoNLL-U Tag Mapping). 3. **Linguistics:** Discourse Planner (Context & Pronominalization). 4. **Automation:** The "Architect" Agent (Generative Grammar Creation). 5. **DevOps:** Interactive QA (Auto-Ticketing & Gold Standard Validation). 6. **Core Optimization:** Weighted Topology Factory (Udiron-style Linearization). 7. **Hybridization:** Learned Micro-Planning (Style Injection). --- ## 2. Component A: The Ninai Bridge (Input Adapter) **Role:** Transforms the Architect into a native renderer for Abstract Wikipedia by accepting `Ninai` JSON Object structures directly. ### 2.1 The Translation Logic Unlike the v1.0 regex approach, v2.0 uses a **Recursive Object Walker**. It treats the input as a Python/JSON tree, respecting the specific constructor keys found in the `Ninai` codebase. * **Location:** `app/adapters/ninai.py` * **Input Schema (Ninai JSON):** ```json { "function": "ninai.constructors.Statement", "args": [ { "function": "ninai.types.Bio" }, { "function": "ninai.constructors.List", "args": [ { "function": "ninai.constructors.Entity", "args": ["Q1"] }, { "function": "ninai.constructors.Entity", "args": ["Q2"] } ]} ] } ``` ### 2.2 Extraction Strategy The adapter traverses the JSON tree and applies specific transformation rules based on the `function` key: | Ninai Constructor Key | Target SKA Field | Logic | | --- | --- | --- | | `ninai.types.Bio` | `frame_type` | Static mapping `"bio"`. | | `ninai.types.Event` | `frame_type` | Static mapping `"event"`. | | `ninai.constructors.List` | `N/A` (Intermediate) | **Recursive Flatten:** Calls `_walk()` on all items in `args[]`, joins results with `", "`. | | `ninai.constructors.Entity` | `subject` / `object` | Extracts the QID string (e.g., `"Q42"`) from `args[0]`. | | `ninai.constructors.Statement` | `N/A` (Root) | Maps `args[0]` to type, `args[1]` to subject, etc. | ### 2.3 Data Flow 1. **Ingest:** `POST /generate` detects `Content-Type: application/json` + `X-Format: ninai`. 2. **Parse:** `NinaiAdapter.parse(payload)` initiates the recursive walk. 3. **Map:** Converts the extracted QIDs and strings into a `BioFrame` or `EventFrame` object. 4. **Validate:** Runs strict Pydantic validation (e.g., ensuring `profession` is present for Bio frames). --- ## 3. Component B: The UD Exporter (Output Adapter) **Role:** Enables evaluation against Universal Dependencies (UD) treebanks by converting internal GF trees into the industry-standard CoNLL-U format. ### 3.1 The "Construction-Time Tagging" Strategy Since we generate text constructively (GF) rather than parsing it (UD), we map the **intent** of the RGL functions to UD tags. * **Location:** `app/core/exporters/ud_mapping.py` * **Frozen Mapping Table:** ```python UD_MAP = { "mkCl": {"arg1": "nsubj", "arg2": "root", "arg3": "obj"}, # Subject, Verb, Object "mkS": {"arg1": "root"}, # Sentence Root "mkNP": {"arg1": "det", "arg2": "head"}, # Det + Noun "mkCN": {"arg1": "amod", "arg2": "head"}, # Adj + Noun "UseN": {"arg1": "head"}, # Bare Noun "AdvNP": {"arg1": "head", "arg2": "nmod"} # Noun + Modifier } ``` ### 3.2 Output Format The API supports `Accept: text/x-conllu`. The output mimics a dependency parse: ```text # text = Marie Curie est une physicienne. # source = SemantikArchitect v2.0 (WikiFre) 1 Marie Marie PROPN _ _ 2 nsubj _ _ 2 Curie Curie PROPN _ _ 3 nsubj _ _ 3 est être VERB _ _ 0 root _ _ ``` --- ## 4. Component C: The Discourse Planner (Context) **Role:** Introduces statefulness to the API. It manages a "Session Context" in Redis to handle multi-sentence coherence (e.g., replacing names with pronouns). ### 4.1 Storage Schema We implement a dedicated Pydantic model for the Redis payload. * **Location:** `app/core/domain/context.py` * **Key:** `awa:session:{uuid}` (TTL: 600s) * **Payload (`SessionContext`):** ```json { "session_id": "a1b2-c3d4", "history_depth": 1, "current_focus": { "label": "Marie Curie", "gender": "f", "qid": "Q7186", "recency": 0 } } ``` ### 4.2 The Pronominalization Logic 1. **Check:** Middleware intercepts the request. Is `X-Session-ID` present? 2. **Fetch:** Retrieve `SessionContext` from Redis. 3. **Compare:** Does `frame.subject_qid` == `context.current_focus.qid`? 4. **Mutate:** * **If Match:** Change `frame.subject` to `"She"` (or language-specific pronoun). * **If Mismatch:** Keep name, update `current_focus` to the new subject. 5. **Save:** Write updated context back to Redis. --- ## 5. Component D: The "Architect" Agent (Automation) **Role:** Completely automates the creation of Tier 3 grammars. It replaces the "Pidgin" templates with AI-generated GF code, validated by the compiler. ### 5.1 The "Architect" Workflow This logic runs inside `builder/orchestrator.py`. 1. **Detection:** Scanner identifies a missing language (e.g., `WikiHau.gf`). 2. **Prompting:** Sends the **Frozen System Prompt** (see Ledger) to Gemini/LLM. * *Prompt Context:* "Write a Concrete Grammar for Hausa (SVO). Use `mkS`, `mkCl`, `mkNP`." 3. **Drafting:** Saves the LLM output to `generated/src/WikiHau.gf`. 4. **Verification:** Runs `gf -batch -c WikiHau.gf`. 5. **Repair Loop ("The Surgeon"):** * If compilation fails, feed the error log back to the LLM: "You used `mkN` but Hausa requires `mkN0`. Fix it." * Max Retries: 3. --- ## 6. Component E: Interactive QA (DevOps) **Role:** Closes the quality loop by validating output against "Gold Standard" data and auto-filing bug reports. ### 6.1 Gold Standard Integration We ingest the **Udiron Test Suite** as ground truth. * **Source:** `data/tests/gold_standard.json` (Migrated from Udiron). * **Logic:** The Judge Agent runs a daily regression test: * *Input:* `tests.json` Intent. * *Output:* SKA Generation. * *Metric:* Levenshtein Distance & Semantic Similarity. ### 6.2 The "Whistleblower" (Auto-Ticketing) If a Tier 1 (High Road) language fails a Gold Standard test: 1. **Trigger:** `similarity_score < 0.8`. 2. **Payload:** Constructs a GitHub Issue Markdown body. 3. **Action:** `POST /repos/{org}/{repo}/issues` using `GITHUB_TOKEN`. * *Title:* `[QA] Regression: {Language} - {FrameType}`. * *Body:* "Expected 'X', got 'Y'. Confidence: Low." --- ## 7. Component F: Shared Configuration (Infrastructure) **Role:** Centralizes all v2.0 environment variables to prevent configuration drift. ### 7.1 Settings Update (`config.py`) ```python class Settings(BaseSettings): # Core APP_ENV: str = "development" GF_LIB_PATH: str = "/app/gf-rgl" # v2.0 Architecture REDIS_URL: str = "redis://redis:6379/0" SESSION_TTL_SEC: int = 600 # DevOps / QA GITHUB_TOKEN: Optional[str] = None REPO_URL: str = "https://github.com/org/repo" AI_MODEL_NAME: str = "gemini-1.5-pro" ``` --- ## 8. Component G: Weighted Topology Factory (Tier 3 Upgrade) **Role:** Adapts **Udiron's Linearization Logic** to solve the "Word Order Problem" for generated grammars. ### 8.1 The Problem The current Factory hardcodes `SVO` (`subj ++ verb ++ obj`). This produces grammatically incorrect output for languages like Japanese (SOV) or Irish (VSO). ### 8.2 The Solution: Topology Weights We introduce a configuration file that defines relative positions for syntactic roles. * **Location:** `data/config/topology_weights.json` * **Schema:** ```json { "SVO": { "nsubj": -10, "root": 0, "obj": 10 }, "SOV": { "nsubj": -10, "obj": -5, "root": 0 }, "VSO": { "root": -10, "nsubj": 0, "obj": 10 } } ``` ### 8.3 Integration Logic In `utils/grammar_factory.py`, the generation logic is refactored: 1. **Lookup:** Check `MISSING_LANGUAGES[lang]['order']` (e.g., "SOV"). 2. **Assign:** Retrieve weights: `subj (-10)`, `obj (-5)`, `verb (0)`. 3. **Sort:** Order the components by weight. 4. **Generate:** Emit the GF `lin` rule in the correct sorted order: * `lin S = mkS (subj ++ obj ++ verb);` --- ## 9. Component H: Learned Micro-Planning (Hybridization) **Role:** Injects stylistic variation using a lightweight AI call *before* the rigorous GF rendering. ### 9.1 The Logic 1. **Input:** API receives `style="formal"`. 2. **Intercept:** `MicroPlanner` passes the frame to LLM: "Rewrite this BioFrame to be formal. Change 'died' to 'passed away'." 3. **Render:** The *modified* frame is passed to the GF engine. 4. **Result:** Grammatically perfect text (GF) with stylistically varied vocabulary (AI). --- ## 10. Implementation Roadmap To execute this "Omni-Upgrade" without breaking the build, follow this strict order: 1. **Phase 1: Foundation** * Update `config.py` (Pydantic settings). * Create `data/config/topology_weights.json`. 2. **Phase 2: Adapters (Deterministic)** * Implement `NinaiAdapter` (Recursive JSON). * Implement `UDMapping` (Table-based). 3. **Phase 3: Core Logic (Optimization)** * Upgrade `grammar_factory.py` with Weighted Topology. * Implement `SessionContext` & Redis hooks. 4. **Phase 4: AI Services (Probabilistic)** * Create `prompts.py` (Frozen Prompts). * Upgrade `Judge` with Gold Standard data & GitHub Client. 5. **Phase 5: Integration** * Wire middleware into `api.py`. * Deploy `builder/orchestrator.py` with the Architect Agent. ================================================================================================ FILE: docs/Technical-Reference/UPGRADE v2-0 Variable and Configuration Ledger.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: a76dcdb9011ccc287f003387b8363393397835b9274c1914e6395b6c90aa1500 CONTENT_BYTES: 8642 ================================================================================================ Here is the final, **Exhaustive v2.0 Variable & Configuration Ledger**. I have added **Section 10 (Topology Weights)** and **Section 11 (Gold Standard Paths)** to capture the Udiron-derived configurations that were missing from the previous draft. This document is now fully aligned with the v2.0 Architecture Spec. You should save this as **`docs/14-VAR_FIX_LEDGER.md`**. --- ### 🔐 v2.0 Variable & Configuration Ledger (Exhaustive) **SemantiK Architect** #### 1. Purpose This document "freezes" all variable names, database keys, and protocol strings for the v2.0 Omni-Upgrade. **All code implementations must copy-paste these values exactly.** Do not invent new variable names or logic paths. --- #### 2. Environment Variables (`.env`) These must be defined in `app/shared/config.py` using Pydantic `BaseSettings`. | Variable Name | Type | Default / Example Value | Description | | --- | --- | --- | --- | | **`APP_ENV`** | `str` | `"development"` | Toggles debug mode (`development`, `staging`, `production`). | | **`LOG_LEVEL`** | `str` | `"INFO"` | Logging verbosity (`DEBUG`, `INFO`, `WARNING`, `ERROR`). | | **`REDIS_URL`** | `str` | `"redis://redis:6379/0"` | Connection string for Session Store (Discourse Context). | | **`SESSION_TTL_SEC`** | `int` | `600` | Expiration time for Discourse Context keys (10 mins). | | **`GITHUB_TOKEN`** | `str` | `""` | PAT (Personal Access Token) for the Judge Agent (QA). | | **`REPO_URL`** | `str` | `"https://github.com/my-org/awa"` | Target repository URL for QA Issue creation. | | **`AI_MODEL_NAME`** | `str` | `"gemini-1.5-pro"` | The specific LLM model version for Architect & Judge agents. | | **`GOOGLE_API_KEY`** | `str` | `""` | API Key for Google Gemini services. | | **`GF_LIB_PATH`** | `str` | `"/app/gf-rgl"` | Absolute path to the RGL source inside the container. | | **`PGF_PATH`** | `str` | `"/app/gf/semantik_architect.pgf"` | Path to the compiled binary grammar file. | --- #### 3. Redis Schema **Namespace:** `awa:` | Key Pattern | Data Structure | TTL | Description | | --- | --- | --- | --- | | **`awa:session:{session_id}`** | `JSON String` | `SESSION_TTL_SEC` | Stores the `SessionContext` object for pronominalization. | | **`awa:lock:build`** | `String` | `60s` | Mutex lock for the Architect Agent to prevent concurrent builds. | | **`awa:cache:lexicon:{lang}`** | `JSON String` | `3600s` | Caches heavy lexicon files to avoid disk I/O on every request. | **JSON Payload (`SessionContext`):** ```json { "session_id": "uuid-v4", "history_depth": 0, "current_focus": { "label": "Marie Curie", "gender": "f", "qid": "Q7186", "recency": 1 } } ``` --- #### 4. Ninai Protocol Constants (JSON/Object) The `NinaiAdapter` must parse JSON object structures, not LISP strings. The keys below map strictly to the `ninai` Python API and internal logic. | Constant Name | Value / Key | Description | | --- | --- | --- | | **`KEY_FUNC`** | `"function"` | JSON key identifying the constructor class name. | | **`KEY_ARGS`** | `"args"` | JSON key containing the list of arguments for the constructor. | | **`CLS_LIST`** | `"ninai.constructors.List"` | Identifies a Ninai List object (trigger for recursive flattening). | | **`CLS_STATEMENT`** | `"ninai.constructors.Statement"` | Identifies a declarative sentence payload. | | **`CLS_SUPPORT`** | `"ninai.constructors.Support"` | Identifies reference/metadata wrappers (ignored or stripped). | | **`CLS_ENTITY`** | `"ninai.constructors.Entity"` | Wrapper for QIDs (e.g., `args=["Q42"]`). | | **`CLS_TYPE_BIO`** | `"ninai.types.Bio"` | Maps to `frame_type="bio"`. | | **`CLS_TYPE_EVENT`** | `"ninai.types.Event"` | Maps to `frame_type="event"`. | --- #### 5. Universal Dependencies (UD) Truth Table This dictionary **MUST** be implemented exactly in `app/core/exporters/ud_mapping.py`. It defines the rigid mapping between RGL functions and UD tags. ```python # FROZEN DICTIONARY UD_MAP = { # Clause Level "mkCl": {"arg1": "nsubj", "arg2": "root", "arg3": "obj"}, "mkS": {"arg1": "root"}, "mkUtt": {"arg1": "root"}, "mkQS": {"arg1": "root"}, # Question Sentence # Noun Phrase Level "mkNP": {"arg1": "det", "arg2": "head"}, "mkCN": {"arg1": "amod", "arg2": "head"}, "UseN": {"arg1": "head"}, "AdvNP": {"arg1": "head", "arg2": "nmod"}, "DetCN": {"arg1": "det", "arg2": "head"}, # Verb Phrase Level "mkVP": {"arg1": "head", "arg2": "obj"}, # Basic V + O "AdvVP": {"arg1": "head", "arg2": "advmod"}, # Fallback "DEFAULT": "dep" } ``` --- #### 6. AI System Prompts (Frozen Strings) The "Architect Agent" must use this **exact** system prompt in `ai_services/prompts.py` to guarantee deterministic, code-only output. **Constant Name:** `ARCHITECT_SYSTEM_PROMPT` > "You are the SemantiK Architect, an expert in Grammatical Framework (GF). Your task is to write a Concrete Grammar file (*.gf) for a specific language. > **CRITICAL RULES:** > 1. Output **ONLY** the raw GF code. > 2. **NO** Markdown code blocks (```). > 3. **NO** conversational filler ('Here is the code...'). > 4. Implement the 'SemantikArchitect' interface exactly. > 5. Use standard RGL modules: `Syntax`, `Paradigms`." > > --- #### 7. Interactive QA Payload The JSON payload sent by the Judge Agent to the GitHub API. **Endpoint:** `POST /repos/{owner}/{repo}/issues` | Field | Value Template | | --- | --- | | **`title`** | `"[QA] Poor Quality: {lang} - {frame_type}"` | | **`labels`** | `["linguistics", "auto-generated", "v2-audit"]` | | **`body`** | See template below. | **Body Template:** ```markdown ### 🚨 Linguistic Quality Alert **Language:** {lang} **Frame:** {frame_type} **Confidence:** {confidence_score} **Generated Output:** > "{generated_text}" **Critique:** {judge_critique} **Metadata:** * Engine: {engine_name} * Strategy: {strategy} (Tier 1/3) * Session ID: {session_id} *Reported by SKA Judge Agent* ``` --- #### 8. Pydantic Model Fields (Domain Objects) To avoid `KeyError` exceptions, these class definitions in `app/core/domain/` are final. ### `DiscourseEntity` (Class) * `label`: `str` (The surface text, e.g., "Marie Curie") * `gender`: `str` (Must be one of: `"m", "f", "n", "c"`) * `qid`: `str` (Must match regex `^Q\d+$`, e.g., "Q42") * `recency`: `int` (0 for current, incremented each turn) ### `SessionContext` (Class) * `session_id`: `str` (UUID4) * `history_depth`: `int` (Default 0) * `current_focus`: `Optional[DiscourseEntity]` (The entity currently in focus for pronominalization) ### `BioFrame` (Update for v2.0) * `frame_type`: `Literal["bio"]` * `name`: `str` * `profession`: `str` (Can be comma-separated list) * `nationality`: `Optional[str]` * `gender`: `Optional[Literal["m", "f", "n"]]` * `context_id`: `Optional[str]` (New field for session linking) --- #### 9. Error Codes (v2.0 Extension) These map to specific HTTP Status Codes and internal Exception classes. | HTTP Code | Exception Class | Description | | --- | --- | --- | | **400** | `NinaiParseError` | Malformed JSON or invalid constructor key in input payload. | | **400** | `FrameValidationError` | Input frame missing required fields (e.g., `name` in BioFrame). | | **404** | `LanguageNotFoundError` | The requested language ISO code is not in the PGF binary. | | **409** | `SessionConflict` | Concurrent write to same Redis session key (Optimistic Locking failure). | | **422** | `LexiconMissingError` | A required word (e.g., profession) is missing from `people.json`. | | **424** | `RGLMappingError` | UD Exporter encountered an unmapped RGL function and `DEFAULT` fallback failed. | | **503** | `AgentQuotaExceeded` | The Architect Agent hit the Gemini/OpenAI API rate limit. | | **500** | `GFRuntimeError` | The C-runtime for PGF crashed or returned null. | --- #### 10. Topology Weights Schema (Udiron Integration) Used by `utils/grammar_factory.py` to order dependencies for Tier 3 languages. **File:** `data/config/topology_weights.json` ```json { "SVO": { "nsubj": -10, "root": 0, "obj": 10 }, "SOV": { "nsubj": -10, "obj": -5, "root": 0 }, "VSO": { "root": -10, "nsubj": 0, "obj": 10 }, "VOS": { "root": -10, "obj": 5, "nsubj": 10 }, "OVS": { "obj": -10, "root": 0, "nsubj": 10 }, "OSV": { "obj": -10, "nsubj": -5, "root": 0 } } ``` --- #### 11. Gold Standard Paths Paths to the validation datasets ingested from Udiron. | Key | Path | Description | | --- | --- | --- | | **`GOLD_TESTS_PATH`** | `data/tests/gold_standard.json` | The `tests.json` file migrated from Udiron. | | **`GOLD_SCHEMA_PATH`** | `data/tests/schema.json` | Validation schema for test cases. | ================================================================================================ FILE: docs/Technical-Reference/UPGRADE_v2.1_SPECIFICATION.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: ab574b1922e657556c99acaa6dae55d6fc0102da767d1ce42c00746e0c7dffeb CONTENT_BYTES: 6600 ================================================================================================ # SemantiK Architect: v2.1 Upgrade Specification **Codename:** "The Brain Transplant" **Status:** IMPLEMENTED **Date:** December 23, 2025 ## 1. Executive Summary This upgrade activates **Zone B (Lexicon)**, transforming the Architect from a system that merely passes strings (`"Douglas Adams"`) to a system that understands concepts (`Q42`). It solves the "Cold Start" problem via lazy loading and resolves the "Triangle of Doom" by linking the Abstract Grammar directly to the `gf-wordnet` repository. **Key Capabilities:** * **Entity Grounding:** Resolves `QIDs` (Q42) into concrete GF functions (`douglas_adams_PN`) using a harvested dictionary of ~380k words. * **Semantic Framing:** Replaces generic triples with specific frames (`mkBio`, `mkEvent`). * **Smart Overloading:** Dynamically selects grammar functions (`mkBioFull` vs `mkBioProf`) based on data availability (P106/P27). --- ## 2. The Data Pipeline (ETL) We have replaced the manual "Copy-Paste" strategy with an automated **Universal Harvester**. ### Component: The Harvester * **Path:** `tools/harvest_lexicon.py` * **Source 1 (The Mine):** Local `gf-wordnet` repo. * *Logic:* Parses `WordNet.gf` (Abstract) and `WordNet{Lang}.gf` (Concrete) to map `02756049-n` `apple_N` `"pomme"`. * *Robustness:* Implements recursive search to find files hidden in subdirectories (e.g., `gf/bul/WordNetBul.gf`). * **Source 2 (The Cloud):** Wikidata (SPARQL). * *Logic:* Fetches labels for generic entities (Names, Cities) missing from WordNet. * **Output:** Generates `data/lexicon/{lang}/wide.json`. * *Format:* JSON Shards optimized for O(1) Python lookup. --- ## 3. The Runtime Architecture (Python) The "Brain" of the system has been rewired to hold the massive lexicon in memory without crashing the worker. ### A. The Memory Bank (`app/shared/lexicon.py`) * **Role:** Singleton In-Memory Database. * **Mechanism:** **Lazy Loading**. It does not load all 77 languages at startup. It loads a language's `wide.json` shard (50MB+) only when a request for that language arrives. * **Lookup Strategy:** 1. **Primary:** QID Match (`Q42` `douglas_adams_PN`). 2. **Fallback:** Lemma Match (`"Douglas Adams"` `douglas_adams_PN`). ### B. The Bridge (`app/adapters/ninai.py`) * **Role:** Converts Ninai JSON GF Abstract Syntax Tree. * **Upgrade:** Integrated `LexiconRuntime`. * **Logic Flow:** 1. Receive `{"function": "mkBio", "args": ["Q42", "Q123"]}`. 2. Ask Lexicon: "What is Q42 in Russian?" Returns `douglas_adams_PN`. 3. Ask Lexicon: "What is Q123 in Russian?" Returns `edinburgh_PN`. 4. Generate Tree: `mkBio (douglas_adams_PN) (edinburgh_PN)`. * **Safety:** If a QID is missing from the lexicon, it degrades gracefully to a String Literal: `mkBio (mkPN "Douglas Adams") ...`. ### C. The Worker (`app/workers/worker.py`) * **Startup Routine:** 1. Loads `Wiki.pgf` (Zone A Grammar). 2. **Pre-loads** `eng` Lexicon (Zone B) to ensure zero latency for the default language. * **Zombie Filter:** Checks `everything_matrix.json` before loading a language. If `verdict.runnable` is `False`, the language is purged from the runtime to prevent user-facing errors. --- ## 4. The Grammar Architecture (GF) We have established the **"Triangle of Doom"** alignment between the Schema (Abstract) and the Vocabulary (Concrete). ### A. The Schema (`gf/semantik_architect.gf`) Defines the strict API contract. To handle real-world data variance, we use **Overloading**: ```haskell cat Statement ; Entity ; Profession ; Nationality ; fun -- The "Perfect" Case (P106 + P27 exist) mkBioFull : Entity -> Profession -> Nationality -> Statement ; -- The "Partial" Cases (Missing Data) mkBioProf : Entity -> Profession -> Statement ; mkBioNat : Entity -> Nationality -> Statement ; -- Type Coercion (Allows WordNet functions to fit our Schema) lexProf : N -> Profession ; lexNat : A -> Nationality ; ``` ### B. The Implementation (`gf/WikiEng.gf`) The Concrete Grammar must now **open** the external WordNet module to access the 380k identifiers. ```haskell concrete WikiEng of SemantikArchitect = open SyntaxEng, ParadigmsEng, WordNetEng in { -- 'open WordNetEng' allows us to use 'physicist_N' directly lin mkBioFull s p n = mkS (mkCl s (mkVP n p)) ; -- "He is an American physicist" mkBioProf s p = mkS (mkCl s (mkVP (mkCN p))) ; -- "He is a physicist" -- Coercion lexProf n = mkCN n ; } ``` --- ## 5. Execution Guide To deploy v2.1, execute the following sequence in **WSL**. ### Phase 1: Harvest (Zone B) ```bash # 1. Harvest local WordNet data (The "Rosetta Stone") python3 tools/harvest_lexicon.py wordnet \ --root "/mnt/c/MyCode/SemantiK_Architect/gf-wordnet" \ --langs eng,rus,bul,swe # 2. Update the Everything Matrix python3 tools/everything_matrix/build_index.py ``` ### Phase 2: Compile (Zone A) ```bash # 1. Define the Schema (with Overloading) echo 'abstract SemantikArchitect = { flags startcat = Statement ; cat Statement ; Entity ; Profession ; Nationality ; fun mkEntity : PN -> Entity ; mkBioFull : Entity -> Profession -> Nationality -> Statement ; mkBioProf : Entity -> Profession -> Statement ; lexProf : N -> Profession ; lexNat : N -> Nationality ; }' > gf/semantik_architect.gf # 2. Define the Concrete (linking WordNet) echo 'concrete WikiEng of SemantikArchitect = open SyntaxEng, ParadigmsEng, WordNetEng in { lincat Statement = S ; Entity = NP ; Profession = CN ; Nationality = AP ; lin mkEntity pn = mkNP pn ; mkBioFull s p n = mkS (mkCl s (mkVP n p)) ; mkBioProf s p = mkS (mkCl s (mkVP p)) ; lexProf n = mkCN n ; lexNat n = mkAP (mkA n) ; }' > gf/WikiEng.gf # 3. Build the PGF (This links everything together) # Note: Use -path to point to the WordNet repo gf -make -output-format=pgf -path ".:/mnt/c/MyCode/SemantiK_Architect/gf-wordnet/gf" gf/WikiEng.gf ``` ### Phase 3: Run ```bash # Start the API Worker python3 -m uvicorn app.main:app --reload ``` --- ## 6. Known Limitations & Next Steps 1. **Memory Footprint:** Loading `wide.json` into Python RAM is temporary. For v3.0, migrate to **Redis** or **SQLite**. 2. **Morphology Gaps:** If `WordNet` provides a Noun (`american_N`) but we need an Adjective (`mkBioNat`), the current grammar attempts coercion. This may fail for languages with complex morphology (e.g., Russian). * *Fix:* Future versions of `NinaiAdapter` should check `pos` tag in `wide.json`. 3. **Fact Harvesting:** Currently, we only harvest *Lexicon* (Words). We need a separate `harvest_facts.py` to fetch P106/P27 claims from Wikidata to populate the `mkBio` arguments automatically. ================================================================================================ FILE: LICENSE_EXCLUSION_RATIONALE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: 5bcd71b8cf1f7e8675abbe609949ab20326befe6e9d72bbf0e3fca1b9f068ab7 CONTENT_BYTES: 2240 ================================================================================================ # Exclusion Rationale (WMF / Denny Vrandečić) This document explains why the SemantiK Architect project license excludes the Wikimedia Foundation (WMF), Denny Vrandečić, and parties related to those entities. ## Summary SemantiK Architect is licensed for broad use, modification, and distribution, except for WMF, Denny Vrandečić, and related parties. The exclusion exists because I do not wish to collaborate with or grant rights to entities and individuals whose conduct, in my view, contradicts the project’s values and the manner in which I expect serious technical work and collaboration to be handled. ## Background I presented SemantiK Architect extensively to WMF and to Denny Vrandečić. The presentations included sufficient information to evaluate the system’s capabilities and to validate that the approach is technically serious and relevant to problems they publicly work on and solicit solutions for. Despite that, the refusal response provided was effectively: “not what we had in mind.” ## Reason for Exclusion In my assessment, that response indicates one of the following: - A lack of technical understanding or skill to properly evaluate the work as presented; or - A dismissal of the work for non-technical reasons, including ego, status, or institutional defensiveness. Either outcome is incompatible with the values this project is intended to represent and promote. ## Values This project is grounded in: - collaboration over gatekeeping, - inclusivity over hierarchy, - the common good over ego or institutional self-protection. When decision-making is driven by dismissal rather than serious evaluation and constructive engagement, I do not consider it aligned with these values. ## Intent This exclusion is not intended to prevent others from using SemantiK Architect, nor to restrict research, learning, or independent collaboration in the broader community. It is specifically intended to withhold rights from WMF, Denny Vrandečić, and related parties due to the above concerns. ## Contact If you believe you are incorrectly categorized as an Excluded Party, or you want to discuss collaboration under a separate written agreement, contact: - [rejean.mccormick@initkoa.org] ================================================================================================ FILE: README.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: dd17124d79552018a8fff1a578230e367a5f44d89e357d70fb14bbbdda26cdd5 CONTENT_BYTES: 9378 ================================================================================================ # Semantik Architect **Industrial-grade NLG system for Abstract Wikipedia, Wikifunctions, and standalone API use.** Semantik Architect is a data-driven Natural Language Generation (NLG) system built around a modular Python backend, GF-based grammar assets, language-specific schema/config data, and a separate Next.js frontend. Instead of maintaining one monolithic renderer per language, the project combines: - **A modular Python backend** for API, orchestration, persistence, and generation - **GF grammars and PGF artifacts** for high-precision generation - **Family-oriented engines and morphology modules** for rule-based text realization - **Frame schemas and structured payloads** for language-agnostic input - **A background worker** for long-running build and onboarding tasks - **A separate frontend** for tools, dashboards, and operator workflows The result is a practical architecture for rule-based NLG that can run as a standalone service while staying aligned with the broader Abstract Wikipedia workflow. --- ## Architecture Overview Semantik Architect is best understood as a **modular monolith** on the Python side, with a **separate frontend application**. ```text SemantiK_Architect/ ├── app/ # Python backend │ ├── core/ # Domain logic, constructions, ports, use cases │ ├── adapters/ # API, worker, engines, persistence, messaging │ └── shared/ # Config, DI container, observability, utilities │ ├── architect_frontend/ # Next.js frontend ├── gf/ # GF source grammars and compiled artifacts ├── schemas/ # JSON schemas for frame payloads ├── tools/ # Diagnostics, audits, QA, indexing, health tools ├── builder/ # Grammar/build orchestration ├── ai_services/ # Optional AI-assisted services └── docs/ # Architecture and operational documentation ```` ### Backend responsibilities * **`app/core/`** contains the domain model, semantic constructions, ports, and use cases * **`app/adapters/`** contains FastAPI routes, generation engines, repositories, and infrastructure adapters * **`app/shared/`** contains configuration, dependency injection, logging, and observability helpers ### Frontend responsibilities * **`architect_frontend/`** is a separate Node/Next.js application * It talks to the backend through the canonical API under **`/api/v1`** * In local development it usually runs on **`:3000`**, while the backend runs on **`:8000`** ### Runtime topology In containerized mode, the stack is split into: * **Redis** * **API backend** * **ARQ worker** * **Next.js frontend** * **Nginx reverse proxy** The reverse proxy exposes the app under: * **UI:** `/semantik_architect/` * **API:** `/semantik_architect/api/v1` --- ## Core Concepts ### 1. Semantic Frames Semantic frames are the language-agnostic inputs to the generator. They describe *what* should be said before any language-specific realization happens. Examples include: * biographical facts * entity descriptions * relational/classification payloads * event payloads Example: ```json { "frame_type": "bio", "subject": { "name": "Marie Curie", "qid": "Q7186" }, "properties": { "profession": "physicist", "nationality": "polish" } } ``` ### 2. Constructions The backend contains reusable constructions for common meaning patterns, such as: * copular classification * transitive events * passive events * topic-comment structures * relative clauses * possession and existential forms These constructions let the system stay semantic-first rather than language-script-first. ### 3. Generation Engines Semantik Architect supports multiple realization strategies: * **GF-backed generation** for higher-precision, grammar-driven output * **Python family/morphology engines** for rule-based realization paths and fallback strategies ### 4. Schemas and Configuration The repository includes a dedicated `schemas/` directory with JSON Schemas for structured frame payloads. These schemas are separate from the Python packaging layer and act as contracts for incoming content and tooling. ### 5. Background Work Long-running operations such as onboarding or building language resources are handled asynchronously through the worker stack rather than blocking the API process. --- ## Quick Start (Docker) The easiest way to run the full stack is Docker Compose. ### Start everything ```bash docker compose up --build ``` ### Main endpoints * **Reverse-proxied UI:** `http://localhost:4000/semantik_architect/` * **Backend docs:** `http://localhost:8000/docs` * **Direct frontend container port:** `http://localhost:3000` * **Redis:** `localhost:6379` ### Health check ```bash curl http://localhost:8000/api/v1/health/ready ``` You can also use the direct health route outside the API prefix: ```bash curl http://localhost:8000/health/ready ``` --- ## API Usage The canonical backend contract lives under **`/api/v1`**. ### Generate text Path-style language selection: ```bash curl -X POST http://localhost:8000/api/v1/generate/eng \ -H "x-api-key: secret" \ -H "Content-Type: application/json" \ -d '{ "frame_type": "bio", "subject": {"name": "Marie Curie"}, "properties": {"profession": "physicist", "nationality": "polish"} }' ``` There is also a payload-driven generation route for clients that send language inside the request body. ### Supported languages ```bash curl http://localhost:8000/api/v1/languages ``` ### Onboard or manage languages Administrative language lifecycle operations live under the management layer and are intended to be protected. Example: ```bash curl -X POST http://localhost:8000/api/v1/languages/ \ -H "x-api-key: secret" \ -H "Content-Type: application/json" \ -d '{"code": "zul", "name": "Zulu", "family": "Bantu"}' ``` --- ## Local Development ### Recommended mental model Keep local development simple: * **one Python environment at the repo root** * **one Node environment in `architect_frontend/`** Do **not** create a separate Python virtual environment for every subdirectory unless you intentionally refactor the repo into multiple installable Python packages. ### Backend setup Run the backend in **WSL/Linux or Docker**. The GF runtime and related dependencies are not a good fit for native Windows execution. ```bash python3 -m venv .venv source .venv/bin/activate pip install -r requirements.txt ``` ### Frontend setup ```bash cd architect_frontend npm install npm run dev ``` ### Run services manually Backend: ```bash source .venv/bin/activate uvicorn app.adapters.api.main:create_app --factory --host 0.0.0.0 --port 8000 --reload ``` Worker: ```bash source .venv/bin/activate arq app.workers.worker.WorkerSettings --watch app ``` Frontend: ```bash cd architect_frontend npm run dev ``` ### Unified orchestration For lifecycle operations such as build, doctor, align, and service startup, **`manage.py`** is the canonical orchestrator. Examples: ```bash python manage.py doctor python manage.py align --force python manage.py build --langs en fr ``` --- ## Testing The test suite is organized into: * **unit** tests for core/domain behavior * **integration** tests for adapters and infrastructure-aware flows * **e2e** tests for API behavior Typical commands: ```bash pytest pytest tests/unit pytest tests/integration ``` --- ## Build and Grammar Layer Semantik Architect includes a GF build/orchestration pipeline and a compiled PGF runtime artifact. Important directories: * `gf/` for grammar sources and compiled output * `builder/` for build orchestration * `tools/` for audits, inventory, diagnostics, and health checks If you are working on the grammar layer, prefer the documented build flow through `manage.py` and the builder/orchestrator tooling rather than ad hoc commands. --- ## Tools and Operator Workflows The backend also exposes a protected **Tools API** used by the frontend dashboard. This layer is designed around: * an allowlisted registry of tool commands * repo-root path confinement * argument allowlisting * output truncation and timeouts * optional AI-gated tools This keeps operational tooling usable from the UI without turning the backend into a generic remote shell. --- ## Status Current repository direction: * canonical FastAPI backend under **`app/adapters/api/main.py`** * canonical API contract under **`/api/v1`** * separate Next.js frontend under **`architect_frontend/`** * async worker for background processing * Dockerized multi-service deployment * schema-driven input contracts * observability hooks for structured logging and tracing --- ## Repository Map * **Backend:** `app/` * **Frontend:** `architect_frontend/` * **Grammars:** `gf/` * **Schemas:** `schemas/` * **Operational tools:** `tools/` * **Build orchestration:** `builder/` * **AI helpers:** `ai_services/` * **Docs:** `docs/` --- ## Links * **Repository:** `README.md`, `docs/`, and the Docker files in this repo are the best current source of truth * **Setup guide:** `docs/00-SETUP_AND_DEPLOYMENT.md` * **Tools inventory:** `docs/17-TOOLS_AND_TESTS_INVENTORY.md` * **API/UI unification notes:** `docs/APIUI Unification Update.txt` ================================================================================================ FILE: RELEASE.md AUTHORITY: reference CONTENT_ROLE: knowledge CONTENT_SHA256: b2927c0425c3b7c20c18e8610524934ec9d34f41f01830347b26412879dc5cf2 CONTENT_BYTES: 515 ================================================================================================ # RGL releases The RGL does not use semantic versioning. Releases are instead made periodically, as snapshots of the current state of the library. Releases are Git tagged `YYYYMMDD`, and for each release a binary package (as `.gfo` files) is made available as a GitHub release. ## Creating a new release 1. Run the "Create release" workflow through the GitHub actions interface (instructions [here](https://docs.github.com/en/free-pro-team@latest/actions/managing-workflow-runs/manually-running-a-workflow)).