crawl4ai

Author	SHA1	Message	Date
UncleCode	a3cb938675	feat(theme): enable dark color mode in mkdocs configuration	2025-05-16 21:44:56 +08:00
UncleCode	9b60988232	feat(feedback): add feedback modal styles and integrate into mkdocs configuration	2025-05-16 21:25:10 +08:00
UncleCode	98e951f611	fix(mkdocs): remove duplicate gtag.js entry in extra_javascript	2025-05-16 20:52:41 +08:00
UncleCode	baca2df8df	feat(analytics): add Google Tag Manager script and gtag.js for tracking	2025-05-16 20:49:02 +08:00
UncleCode	8a5e23d374	feat(crawler): add separate timeout for wait_for condition Adds a new wait_for_timeout parameter to CrawlerRunConfig that allows specifying a separate timeout for the wait_for condition, independent of the page_timeout. This provides more granular control over waiting behaviors in the crawler. Also removes unused colorama dependency and updates LinkedIn crawler example. BREAKING CHANGE: LinkedIn crawler example now uses different wait_for_images timing	2025-05-16 17:00:45 +08:00
ntohidi	22725ca87b	fix(crawler): initialize `captured_console` to prevent unbound local error for local HTML files. REF: #1072 Resolved a bug where running the crawler on local HTML files with `capture_console_messages=False` (default) raised `UnboundLocalError` due to `captured_console` being accessed before assignment.	2025-05-15 11:29:36 +02:00
ntohidi	e0fbd2b0a0	fix(schema): update `f` parameter description to use lowercase enum values. REF: #1070 Revised the description for the `f` parameter in the `/mcp/md` tool schema to use lowercase enum values (`raw`, `fit`, `bm25`, `llm`) for consistency with the actual `enum` definition. This change prevents LLM-based clients (e.g., Gemini via LibreChat) from generating uppercase values like `"FIT"`, which caused 422 validation errors due to strict case-sensitive matching.	2025-05-15 10:45:23 +02:00
ntohidi	32966bea11	fix(extraction): resolve `'str' object has no attribute 'choices'` error in LLMExtractionStrategy. Refs: #979 This patch ensures consistent handling of `response.choices[0].message.content` by avoiding redefinition of the `response` variable, which caused downstream exceptions during error handling.	2025-05-15 10:09:19 +02:00
Ahmed-Tawfik94	a3b0cab52a	#1088 is sloved flag -bc now if for --byPass-cache	2025-05-15 11:25:06 +08:00
medo94my	137556b3dc	fix the EXTRACT to match the styling of the other methods	2025-05-14 16:01:10 +08:00
ntohidi	260e2dc347	fix(browser): create browser config before launching managed browser instance. REF: https://discord.com/channels/1278297938551902308/1278298697540567132/1371683009459392716	2025-05-13 14:03:20 +02:00
ntohidi	25d97d56e4	fix(dependencies): remove duplicated aiofiles from project dependencies. REF #1045	2025-05-13 13:56:12 +02:00
Aravind Karnam	98a56e6e01	Merge next branch	2025-05-13 17:12:11 +05:30
Emmanuel Ferdman	1e1c887a2f	fix(docker-api): migrate to modern datetime library API Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>	2025-05-13 00:04:58 -07:00
UncleCode	897e017361	Set version to 0.6.3 vr0.6.3 v0.6.3	2025-05-12 21:20:10 +08:00
UncleCode	a3e9ef91ad	fix(crawler): remove automatic page closure in screenshot methods Removes automatic page closure in take_screenshot and take_screenshot_naive methods to prevent premature closure of pages that might still be needed in the calling context. This allows for more flexible page lifecycle management by the caller. BREAKING CHANGE: Page objects are no longer automatically closed after taking screenshots. Callers must explicitly handle page closure when appropriate.	2025-05-12 21:17:57 +08:00
UncleCode	76dd86d1b3	Merge remote-tracking branch 'origin/linkedin-prep' into next	2025-05-08 17:13:59 +08:00
UncleCode	206a9dfabd	feat(crawler): add session management and view-source support Add session_id feature to allow reusing browser pages across multiple crawls. Add support for view-source: protocol in URL handling. Fix browser config reference and string formatting issues. Update examples to demonstrate new session management features. BREAKING CHANGE: Browser page handling now persists when using session_id	2025-05-08 17:13:35 +08:00
ntohidi	1af3d1c2e0	Merge branch '2025-APR-1' of https://github.com/unclecode/crawl4ai into 2025-APR-1	2025-05-08 11:11:32 +02:00
Aravind Karnam	c1041b9bbe	fix: exclude_external_images flag simply discards elements ref:https://github.com/unclecode/crawl4ai/issues/345	2025-05-07 18:43:29 +05:30
Aravind Karnam	f6e25e2a6b	fix: check_robots_txt to support wildcard rules ref: #699	2025-05-07 17:53:30 +05:30
ntohidi	ee93acbd06	fix(async_playwright_crawler): use config directly instead of self.config for verbosity check	2025-05-07 12:32:38 +02:00
Aravind Karnam	2b17f234f8	docs: update direct passing of content_filter to CrawlerRunConfig and instead pass it via MarkdownGenerator. Ref: #603	2025-05-07 15:20:36 +05:30
ntohidi	eebb8c84f0	fix(requirements): add PyPDF2 dependency for PDF processing	2025-05-07 11:18:44 +02:00
ntohidi	12783fabda	fix(dependencies): update pillow version constraint to allow newer releases. ref #709	2025-05-07 11:18:13 +02:00
Aravind Karnam	39e3b792a1	Merge branch 'next' into 2025-APR-1	2025-05-07 10:25:25 +05:30
Aravind Karnam	aaf05910eb	fix: removed unnecessary imports and installs	2025-05-06 15:53:55 +05:30
Aravind Karnam	a0555d5fa6	merge:from next branch	2025-05-06 15:16:47 +05:30
Aravind Karnam	38ebcbb304	fix: provide support for local llm by adding it to the arguments	2025-05-05 10:34:38 +05:30
UncleCode	9b5ccac76e	feat(extraction): add RegexExtractionStrategy for pattern-based extraction Add new RegexExtractionStrategy for fast, zero-LLM extraction of common data types: - Built-in patterns for emails, URLs, phones, dates, and more - Support for custom regex patterns - LLM-assisted pattern generation utility - Optimized HTML preprocessing with fit_html field - Enhanced network response body capture Breaking changes: None	2025-05-02 21:15:24 +08:00
Aravind Karnam	87d4b0fff4	format bash scripts properly so copy & paste may work without issues	2025-05-02 17:21:09 +05:30
Aravind Karnam	bd5a9ac632	updated readme with arguments for litellm	2025-05-02 17:04:42 +05:30
Aravind Karnam	6650b2f34a	fix: replace openAI with litellm to support multiple llm providers	2025-05-02 16:51:15 +05:30
Aravind Karnam	5cc58f9bb3	fix: 1. duplicate verbose flag 2.inconsistency in argument name --profile-name 3. duplicate initialisaiton of env_defaults	2025-05-02 16:40:58 +05:30
Aravind Karnam	baf7f6a6f5	fix: typo in readme	2025-05-02 16:33:11 +05:30
ntohidi	e0cd3e10de	fix(crawler): initialize captured_console variable for local file processing	2025-05-02 10:35:35 +02:00
UncleCode	94e9959fe0	feat(docker-api): add job-based polling endpoints for crawl and LLM tasks Implements new asynchronous endpoints for handling long-running crawl and LLM tasks: - POST /crawl/job and GET /crawl/job/{task_id} for crawl operations - POST /llm/job and GET /llm/job/{task_id} for LLM operations - Added Redis-based task management with configurable TTL - Moved schema definitions to dedicated schemas.py - Added example polling client demo_docker_polling.py This change allows clients to handle long-running operations asynchronously through a polling pattern rather than holding connections open.	2025-05-01 21:24:52 +08:00
Aravind Karnam	7c2fd5202e	fix: incorrect params and commands in linkedin app readme	2025-05-01 18:27:03 +05:30
UncleCode	ee01b81f3e	Merge branch 'merge-pr971' into next	2025-05-01 18:58:41 +08:00
UncleCode	0e5d672763	Merge branch 'pr-971' into merge-pr971	2025-05-01 18:57:28 +08:00
wakaka6	cd2b490b40	refactor(logger): Apply the Enumeration for color	2025-05-01 17:04:44 +08:00
UncleCode	50f0b83fcd	feat(linkedin): add prospect-wizard app with scraping and visualization Add new LinkedIn prospect discovery tool with three main components: - c4ai_discover.py for company and people scraping - c4ai_insights.py for org chart and decision maker analysis - Interactive graph visualization with company/people exploration Features include: - Configurable LinkedIn search and scraping - Org chart generation with decision maker scoring - Interactive network graph visualization - Company similarity analysis - Chat interface for data exploration Requires: crawl4ai, openai, sentence-transformers, networkx	2025-04-30 19:38:25 +08:00
ntohidi	1d6a2b9979	fix(crawler): surface real redirect status codes and keep redirect chain. the 30x response instead of always returning 200. Refs #660	2025-04-30 12:29:17 +02:00
ntohidi	039be1b1ce	feat: add pdf2image dependency to requirements	2025-04-30 11:41:35 +02:00
UncleCode	9499164d3c	feat(browser): improve browser profile management and cleanup Enhance browser profile handling with better process cleanup and documentation: - Add process cleanup for existing Chromium instances on Windows/Unix - Fix profile creation by passing complete browser config - Add comprehensive documentation for browser and CLI components - Add initial profile creation test - Bump version to 0.6.3 This change improves reliability when managing browser profiles and provides better documentation for developers.	2025-04-29 23:04:32 +08:00
Marc Sacristán	53245e4e0e	Fix: README.md urls list	2025-04-29 16:26:35 +02:00
UncleCode	2140d9aca4	fix(browser): correct headless mode default behavior Modify BrowserConfig to respect explicit headless parameter setting instead of forcing True. Update version to 0.6.2 and clean up code formatting in examples. BREAKING CHANGE: BrowserConfig no longer defaults to headless=True when explicitly set to False	2025-04-26 21:09:50 +08:00
UncleCode	ccec40ed17	feat(models): add dedicated tables field to CrawlResult - Add tables field to CrawlResult model while maintaining backward compatibility - Update async_webcrawler.py to extract tables from media and pass to tables field - Update crypto_analysis_example.py to use the new tables field - Add /config/dump examples to demo_docker_api.py - Bump version to 0.6.1	2025-04-24 18:36:25 +08:00
Aravind Karnam	094201ab2a	Merge next + resolve conflicts	2025-04-23 19:44:50 +05:30
UncleCode	ad4dfb21e1	Remoce "rc1"	2025-04-23 21:00:00 +08:00

... 3 4 5 6 7 ...

1045 Commits