archon

mirror of https://github.com/coleam00/Archon.git synced 2025-12-24 02:39:17 -05:00

Author	SHA1	Message	Date
leex279	8777e9456c	feat: Prioritize same-directory discovery for llms.txt and sitemaps Improve discovery logic to check the same directory as the base URL first before falling back to root-level and subdirectories. This ensures files like https://supabase.com/docs/llms.txt are found when crawling https://supabase.com/docs. Changes: - Check same directory as base_url first (e.g., /docs/llms.txt for /docs URL) - Fall back to root-level urljoin behavior - Include base directory name in subdirectory checks (e.g., /docs subdirectory) - Maintain priority order: same-dir > root > subdirectories - Log discovery location for better debugging This addresses cases where documentation directories contain their own llms.txt or sitemap files that should take precedence over root-level files. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-17 19:26:24 +02:00
leex279	e5160dde5c	fix: Address CodeRabbit feedback for discovery service - Preserve URL case in robots.txt parsing by only lowercasing the sitemap: prefix check - Add support for relative sitemap paths in robots.txt using urljoin() - Fix HTML meta tag parsing to use case-insensitive regex instead of lowercasing content - Add URL scheme validation for discovered sitemaps (http/https only) - Fix discovery target domain filtering to use discovered URL's domain instead of input URL - Clean up whitespace and improve dict comprehension usage These changes improve discovery reliability and prevent URL corruption while maintaining backward compatibility with existing discovery behavior. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-17 19:03:25 +02:00
leex279	968e5b73fe	Add SSL verification and response size limits to discovery service - Enable SSL certificate verification (verify=True) for all HTTP requests - Implement streaming with size limits (10MB default) to prevent memory exhaustion - Add _read_response_with_limit() helper for secure response reading - Update all test mocks to support streaming API with iter_content() - Fix test assertions to expect new security parameters - Enforce deterministic rounding in progress mapper tests Security improvements: - Prevents MITM attacks through SSL verification - Guards against DoS via oversized responses - Ensures proper resource cleanup with response.close() 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-14 22:31:19 +02:00
leex279	d696918ff0	Merge main into feature/automatic-discovery-llms-sitemap-430 Resolved merge conflicts by integrating features from both branches: - Added page_storage_ops service initialization from main - Merged link text extraction with discovery mode features - Preserved discovery single-file mode and domain filtering - Maintained link text fallbacks for title extraction 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-11 09:31:24 +02:00
sean-eskerium	4050c3540a	Merge pull request #777 from coleam00/refactor/projects-ui Refactor the UI and add Documents back.	2025-10-10 21:58:25 -04:00
Developer	ef4262681f	Code rabbit issues fix again	2025-10-10 21:54:04 -04:00
Cole Medin	77e9342c27	Updating title exxtraction for llms.txt	2025-10-10 18:16:03 -05:00
Cole Medin	4a9ed51cff	Adjusting table creation order in complete_setup.sql	2025-10-10 17:55:30 -05:00
Cole Medin	571e7c18c4	Correcting migrations in complete_setup.sql	2025-10-10 17:52:14 -05:00
Cole Medin	710909eecd	Fixing up migration order	2025-10-10 17:50:41 -05:00
Developer	913f47ba62	code rabbit feedback	2025-10-10 18:40:25 -04:00
Developer	20c57acb00	Code rabbit feedback	2025-10-10 18:30:12 -04:00
DIY Smart Code	3168c8b69f	fix: Set explicit PLAYWRIGHT_BROWSERS_PATH to fix browser installation (#765 ) * fix: Set explicit PLAYWRIGHT_BROWSERS_PATH to fix browser installation Fixes Playwright browser not found error during web crawling. The issue was introduced in the uv migration (`9f22659`) where the browser installation path was not explicitly set as a persistent environment variable. Changes: - Add ENV PLAYWRIGHT_BROWSERS_PATH=/ms-playwright - Add --with-deps flag to playwright install command - Add comprehensive root cause analysis document Without this fix, Playwright installed browsers to a default location at build time but couldn't find them at runtime, causing crawling operations to fail with "Executable doesn't exist" errors. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: Remove --with-deps flag to prevent build conflicts The --with-deps flag was causing build failures on some systems because: - We already manually install all Playwright dependencies (lines 26-49) - --with-deps attempts to reinstall these packages - This causes package conflicts and build failures on Windows/WSL The core fix (ENV PLAYWRIGHT_BROWSERS_PATH) remains the same. * Delete PLAYWRIGHT_FIX_ANALYSIS.md --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Cole Medin <cole@dynamous.ai>	2025-10-10 17:11:52 -05:00
sean-eskerium	7c3823e08f	Fixes: crawl code storage issue with <think> tags for ollama models. (#775 ) * Fixes: crawl code storage issue with <think> tags for ollama models. * updates from code rabbit review	2025-10-10 17:09:53 -05:00
Developer	8ff39fa1d5	Merge branch 'main' into refactor/projects-ui Merged in PR #776 (refactor/knowledge-ui) from main. No conflicts - different features.	2025-10-10 17:08:05 -04:00
sean-eskerium	94e28f85fd	Merge pull request #776 from coleam00/refactor/knowledge-ui Refactoring the UI for consistent styling	2025-10-10 17:03:05 -04:00
sean-eskerium	e22c6c3836	fix code rabbit suggestions.	2025-10-10 14:42:01 -04:00
sean-eskerium	a860b27848	Refactor the UI and add Documents back.	2025-10-10 14:24:09 -04:00
sean-eskerium	691adccc12	Refactoring the UI for consistent styling	2025-10-10 03:36:35 -04:00
sean-eskerium	4ad1fb0808	Merge pull request #772 from coleam00/feature/ui-style-guide Feature/UI style guide	2025-10-09 21:21:17 -04:00
sean-eskerium	88cb8d7f03	Update archon-ui-main/src/features/style-guide/layouts/ProjectsLayoutExample.tsx Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>	2025-10-09 21:17:00 -04:00
sean-eskerium	f0030699a8	Update archon-ui-main/src/features/style-guide/layouts/ProjectsLayoutExample.tsx Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>	2025-10-09 21:15:14 -04:00
sean-eskerium	59f4568fda	another round of code rabbit feedback	2025-10-09 21:05:12 -04:00
sean-eskerium	0013336ee3	Merge branch 'main' of https://github.com/coleam00/Archon into feature/ui-style-guide	2025-10-09 20:42:08 -04:00
sean-eskerium	ad82f6e9f6	Another round of Coderabbit feedback.	2025-10-09 20:40:47 -04:00
Cole Medin	bfd0a84f64	RAG Enhancements (Page Level Retrieval) (#767 ) * Initial commit for RAG by document * Phase 2 * Adding migrations * Fixing page IDs for chunk metadata * Fixing unit tests, adding tool to list pages for source * Fixing page storage upsert issues * Max file length for retrieval * Fixing title issue * Fixing tests	2025-10-09 19:39:27 -05:00
sean-eskerium	c3f42504ea	code rabbit updates	2025-10-09 20:19:51 -04:00
sean-eskerium	98946817b4	Merge remote-tracking branch 'origin/main' into feature/ui-style-guide	2025-10-09 17:43:43 -04:00
sean-eskerium	02533dc37c	Fixing Code Rabbit suggestions.	2025-10-09 16:23:32 -04:00
DIY Smart Code	e6d538fdd8	Merge pull request #769 from coleam00/crawl4ai-update chore: update crawl4ai from 0.6.2 to 0.7.4	2025-10-09 21:52:36 +02:00
sean-eskerium	daf915c083	Fixes from biome and consistency review.	2025-10-09 14:26:37 -04:00
sean-eskerium	4e6116fa2f	Fix consistency and biome formatting issues	2025-10-09 13:49:12 -04:00
sean-eskerium	9e4c7eaf4e	Updating documentation and the review command refinement.	2025-10-09 13:35:24 -04:00
sean-eskerium	db538a5f46	Remove dead code	2025-10-09 12:14:36 -04:00
sean-eskerium	5c7924f43d	Merge main into feature/ui-style-guide - Resolved package-lock.json conflict - Kept Tailwind 4.1.2 upgrade from feature branch - Merged main's updates (react-icons, file reorganization, new features) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-09 11:53:27 -04:00
sean-eskerium	6f173e403d	remove prp docs	2025-10-09 11:49:41 -04:00
sean-eskerium	bebe4c1037	candidate for release	2025-10-09 11:49:03 -04:00
Wirasm	489415d723	Fix: Database timeout when deleting large sources (#737 ) * fix: implement CASCADE DELETE for source deletion timeout issue - Add migration 009 to add CASCADE DELETE constraints to foreign keys - Simplify delete_source() to only delete parent record - Database now handles cascading deletes efficiently - Fixes timeout issues when deleting sources with thousands of pages * chore: update complete_setup.sql to include CASCADE DELETE constraints - Add ON DELETE CASCADE to foreign keys in initial setup - Include migration 009 in the migrations tracking - Ensures new installations have CASCADE DELETE from the start	2025-10-09 17:52:06 +03:00
DIY Smart Code	00fe2599ad	Delete python/test_url_resolution_fix.py	2025-10-09 16:05:37 +02:00
DIY Smart Code	f9a506b9c9	Delete CRAWL4AI_UPDATE.md	2025-10-09 16:04:58 +02:00
sean-eskerium	2e68403db0	update styles of the primitives.	2025-10-09 09:51:50 -04:00
sean-eskerium	80992ca975	Epgrade to Tailwind 4	2025-10-09 09:31:47 -04:00
sean-eskerium	70b6e70a95	trying to make the ui reviews programmatic	2025-10-09 07:59:54 -04:00
sean-eskerium	4cb7c46d6e	fixing document browser and updating primitive tab styles.	2025-10-09 00:15:29 -04:00
sean-eskerium	17ca62ceb4	refining	2025-10-08 23:43:43 -04:00
sean-eskerium	5b839a1465	command for UI review, and settings to use primitives.	2025-10-08 18:38:12 -04:00
sean-eskerium	0727245c9d	Udate the projects layout. And style guide.	2025-10-08 17:37:29 -04:00
leex279	8deee6fd7a	chore: update crawl4ai from 0.6.2 to 0.7.4 Updates crawl4ai dependency to latest stable version with performance and stability improvements. Key improvements in 0.7.4: - LLM-powered table extraction with intelligent chunking - Fixed dispatcher bug for better concurrent processing - Resolved browser manager race conditions - Enhanced URL processing and proxy support All existing tests pass (18/18). No breaking changes identified. API remains backward compatible. ⚠️ IMPORTANT: URL Resolution Bug Status A critical bug in v0.6.2 where ../../ paths only go up ONE directory instead of TWO has been documented (see crawler-test branch). Status in v0.7.4 is UNKNOWN - testing required before production deployment. Test script provided: python/test_url_resolution_fix.py Related issues fixed in v0.7.x: - #570: General relative URL handling - #1268: URLs after redirects - #1323: Trailing slash base URL handling 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>	2025-10-08 22:27:15 +02:00
sean-eskerium	6e86fd0d9b	updates to style guide components	2025-10-08 13:50:04 -04:00
Josh	a580fdfe66	Feature/LLM-Providers-UI-Polished (#736 ) * Add Anthropic and Grok provider support * feat: Add crucial GPT-5 and reasoning model support for OpenRouter - Add requires_max_completion_tokens() function for GPT-5, o1, o3, Grok-3 series - Add prepare_chat_completion_params() for reasoning model compatibility - Implement max_tokens → max_completion_tokens conversion for reasoning models - Add temperature handling for reasoning models (must be 1.0 default) - Enhanced provider validation and API key security in provider endpoints - Streamlined retry logic (3→2 attempts) for faster issue detection - Add failure tracking and circuit breaker analysis for debugging - Support OpenRouter format detection (openai/gpt-5-nano, openai/o1-mini) - Improved Grok provider empty response handling with structured fallbacks - Enhanced contextual embedding with provider-aware model selection Core provider functionality: - OpenRouter, Grok, Anthropic provider support with full embedding integration - Provider-specific model defaults and validation - Secure API connectivity testing endpoints - Provider context passing for code generation workflows 🤖 Generated with [Claude Code](https://claude.ai/code) Co-Authored-By: Claude <noreply@anthropic.com> * fully working model providers, addressing securtiy and code related concerns, throughly hardening our code * added multiprovider support, embeddings model support, cleaned the pr, need to fix health check, asnyico tasks errors, and contextual embeddings error * fixed contextual embeddings issue * - Added inspect-aware shutdown handling so get_llm_client always closes the underlying AsyncOpenAI / httpx.AsyncClient while the loop is still alive, with defensive logging if shutdown happens late (python/src/server/services/llm_provider_service.py:14, python/src/server/ services/llm_provider_service.py:520). * - Restructured get_llm_client so client creation and usage live in separate try/finally blocks; fallback clients now close without logging spurious Error creating LLM client when downstream code raises (python/src/server/services/llm_provider_service.py:335-556). - Close logic now sanitizes provider names consistently and awaits whichever aclose/close coroutine the SDK exposes, keeping the loop shut down cleanly (python/src/server/services/llm_provider_service.py:530-559). Robust JSON Parsing - Added _extract_json_payload to strip code fences / extra text returned by Ollama before json.loads runs, averting the markdown-induced decode errors you saw in logs (python/src/server/services/storage/code_storage_service.py:40-63). - Swapped the direct parse call for the sanitized payload and emit a debug preview when cleanup alters the content (python/src/server/ services/storage/code_storage_service.py:858-864). * added provider connection support * added provider api key not being configured warning * Updated get_llm_client so missing OpenAI keys automatically fall back to Ollama (matching existing tests) and so unsupported providers still raise the legacy ValueError the suite expects. The fallback now reuses _get_optimal_ollama_instance and rethrows ValueError(OpenAI API key not found and Ollama fallback failed) when it cant connect. Adjusted test_code_extraction_source_id.py to accept the new optional argument on the mocked extractor (and confirm its None when present). * Resolved a few needed code rabbit suggestion - Updated the knowledge API key validation to call create_embedding with the provider argument and removed the hard-coded OpenAI fallback (python/src/server/api_routes/knowledge_api.py). - Broadened embedding provider detection so prefixed OpenRouter/OpenAI model names route through the correct client (python/src/server/ services/embeddings/embedding_service.py, python/src/server/services/llm_provider_service.py). - Removed the duplicate helper definitions from llm_provider_service.py, eliminating the stray docstring that was causing the import-time syntax error. * updated via code rabbit PR review, code rabbit in my IDE found no issues and no nitpicks with the updates! what was done: Credential service now persists the provider under the uppercase key LLM_PROVIDER, matching the read path (no new EMBEDDING_PROVIDER usage introduced). Embedding batch creation stops inserting blank strings, logging failures and skipping invalid items before they ever hit the provider (python/src/server/services/embeddings/embedding_service.py). Contextual embedding prompts use real newline characters everywhereboth when constructing the batch prompt and when parsing the models response (python/src/server/services/embeddings/contextual_embedding_service.py). Embedding provider routing already recognizes OpenRouter-prefixed OpenAI models via is_openai_embedding_model; no further change needed there. Embedding insertion now skips unsupported vector dimensions instead of forcing them into the 1536-column, and the backoff loop uses await asyncio.sleep so we no longer block the event loop (python/src/server/services/storage/code_storage_service.py). RAG settings props were extended to include LLM_INSTANCE_NAME and OLLAMA_EMBEDDING_INSTANCE_NAME, and the debug log no longer prints API-key prefixes (the rest of the TanStack refactor/EMBEDDING_PROVIDER support remains deferred). * test fix * enhanced Openrouters parsing logic to automatically detect reasoning models and parse regardless of json output or not. this commit creates a robust way for archons parsing to work throughly with openrouter automatically, regardless of the model youre using, to ensure proper functionality with out breaking any generation capabilities! * updated ui llm interface, added seprate embeddings provider, made the system fully capabale of mix and matching llm providers (local and non local) for chat & embeddings. updated the ragsettings.tsx ui mainly, along with core functionality * added warning labels and updated ollama health checks * ready for review, fixed som error warnings and consildated ollama status health checks * fixed FAILED test_async_embedding_service.py * code rabbit fixes * Separated the code-summary LLM provider from the embedding provider, so code example storage now forwards a dedicated embedding provider override end-to-end without hijacking the embedding pipeline. this fixes code rabbits (Preserve provider override in create_embeddings_batch) suggesting * - Swapped API credential storage to booleans so decrypted keys never sit in React state (archon-ui-main/src/components/ settings/RAGSettings.tsx). - Normalized Ollama instance URLs and gated the metrics effect on real state changes to avoid mis-counts and duplicate fetches (RAGSettings.tsx). - Tightened crawl progress scaling and indented-block parsing to handle min_length=None safely (python/src/server/ services/crawling/code_extraction_service.py:160, python/src/server/services/crawling/code_extraction_service.py:911). - Added provider-agnostic embedding rate-limit retries so Google and friends back off gracefully (python/src/server/ services/embeddings/embedding_service.py:427). - Made the orchestration registry async + thread-safe and updated every caller to await it (python/src/server/services/ crawling/crawling_service.py:34, python/src/server/api_routes/knowledge_api.py:1291). * Update RAGSettings.tsx - header for 'LLM Settings' is now 'LLM Provider Settings' * (RAG Settings) - Ollama Health Checks & Metrics - Added a 10-second timeout to the health fetch so it doesn't hang. - Adjusted logic so metric refreshes run for embedding-only Ollama setups too. - Initial page load now checks Ollama if either chat or embedding provider uses it. - Metrics and alerts now respect which provider (chat/embedding) is currently selected. - Provider Sync & Alerts - Fixed a sync bug so the very first provider change updates settings as expected. - Alerts now track the active provider (chat vs embedding) rather than only the LLM provider. - Warnings about missing credentials now skip whichever provider is currently selected. - Modals & Types - Normalize URLs before handing them to selection modals to keep consistent data. - Strengthened helper function types (getDisplayedChatModel, getModelPlaceholder, etc.). (Crawling Service) - Made the orchestration registry lock lazy-initialized to avoid issues in Python 3.12 and wrapped registry commands (register, unregister) in async calls. This keeps things thread-safe even during concurrent crawling and cancellation. * - migration/complete_setup.sql:101 seeds Google/OpenRouter/Anthropic/Grok API key rows so fresh databases expose every provider by default. - migration/0.1.0/009_add_provider_placeholders.sql:1 backfills the same rows for existing Supabase instances and records the migration. - archon-ui-main/src/components/settings/RAGSettings.tsx:121 introduces a shared credentialprovider map, reloadApiCredentials runs through all five providers, and the status poller includes the new keys. - archon-ui-main/src/components/settings/RAGSettings.tsx:353 subscribes to the archon:credentials-updated browser event so adding/removing a key immediately refetches credential status and pings the corresponding connectivity test. - archon-ui-main/src/components/settings/RAGSettings.tsx:926 now treats missing Anthropic/OpenRouter/Grok keys as missing, preventing stale connected badges when a key is removed. * - archon-ui-main/src/components/settings/RAGSettings.tsx:90 adds a simple display-name map and reuses one red alert style. - archon-ui-main/src/components/settings/RAGSettings.tsx:1016 now shows exactly one red banner when the active provider - Removed the old duplicate Missing API Key Configuration block, so the panel no longer stacks two warnings. * Update credentialsService.ts default model * updated the google embedding adapter for multi dimensional rag querying * thought this micro fix in the google embedding pushed with the embedding update the other day, it didnt. pushing now --------- Co-authored-by: Chillbruhhh <joshchesser97@gmail.com> Co-authored-by: Claude <noreply@anthropic.com>	2025-10-05 13:49:09 -05:00

1 2 3 4 5 ...

279 Commits