14 Commits
Author SHA1 Message Date
saidsurucu f1d3b60efb feat: Add semantic search documentation and bump version to 0.2.0
- Add OpenRouter API configuration guide for Claude Desktop, 5ire, Gemini CLI
- Document semantic search workflow (initial_keyword + query)
- Update tool count to reflect optional semantic search tool
- Bump version to 0.2.0 for semantic search feature release
2025-12-13 19:51:13 +03:00
saidsurucu 93e64bc1fc docs: Improve search_bedesten_semantic parameter descriptions for LLM usage 2025-12-13 19:44:59 +03:00
saidsurucu 77e2748ade feat(semantic-search): Replace local embedding model with OpenRouter API
- Replace EmbeddingGemma local model with OpenRouter API integration
- Use google/gemini-embedding-001 model via OpenRouter (3072 dimensions)
- Add conditional tool registration: auto-disable if OPENROUTER_API_KEY not set
- Add openai and numpy dependencies to pyproject.toml
- Update .env.example with OPENROUTER_API_KEY configuration
- Fix ruff lint issues in semantic_search module
2025-12-13 18:33:47 +03:00
saidsurucuandClaude da146cf3ec feat: Add semantic search tool (search_bedesten_semantic)
Add semantic search capabilities to MCP server:
- Import semantic_search module components
- Add search_bedesten_semantic tool with EmbeddingGemma integration
- Supports intelligent re-ranking of legal decisions
- 5-step process: keyword search → fetch docs → embed → vector search → format

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-12-13 17:20:27 +03:00
saidsurucuandClaude e771c5b3c5 feat: Add semantic search module
Add semantic search capabilities with:
- embedder.py: Text embedding operations
- processor.py: Document processing
- vector_store.py: Vector storage and retrieval

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-12-13 17:15:12 +03:00
saidsurucu 1223b37adb refactor: Rename KİK document parameter to gundemMaddesiId 2025-12-04 15:58:25 +03:00
saidsurucu e26f09aced refactor: Remove Playwright dependency completely from project
- Replace Playwright base image with python:3.12-slim in Dockerfile
- Remove playwright from pyproject.toml dependencies
- Remove ensure_playwright_browsers() function from mcp_server_main.py
- Delete KİK v1 client files (client.py, models.py) - v2 uses httpx
- Delete postinstall.sh Playwright installation script
- Delete obsolete setup.py and requirements.txt.bak
- Remove saidsurucu-yargi-mcp-f5fa007 snapshot directory
- Update client_v2.py docstring to reflect httpx usage
- Regenerate uv.lock without playwright

KİK v2 now uses pure httpx for all HTTP operations with SSL legacy support.
2025-12-04 15:50:20 +03:00
saidsurucu ae5bae2f4a refactor(kik): Replace Playwright with httpx for document retrieval
- Remove Playwright dependency from KİK v2 client
- Use httpx with legacy SSL context for document fetching
- Remove unused imports (requests, base64, subprocess, shutil)
- Simpler and faster implementation
- Tested: 52,856 chars retrieved successfully
2025-12-04 15:43:22 +03:00
saidsurucu a2b50951e9 feat(kik): Implement document ID encryption for KİK v2 API
Reverse engineered the AES-256-CBC encryption used by KİK's Angular web
application to generate document URL hashes from numeric IDs.

Key findings:
- Algorithm: AES-256-CBC with PKCS7 padding
- Key location: ekapv2.kik.gov.tr module 21554 (environment config)
- Output format: IV (16 bytes hex) + Ciphertext (16 bytes hex) = 64 chars

Changes:
- Added encrypt_document_id() static method to KikV2ApiClient
- Updated get_document_markdown() to auto-encrypt numeric gundemMaddesiId
- Added cryptography>=44.0.0 dependency for AES encryption
- Both primary and fallback URL paths now support encryption

This enables direct document retrieval from numeric search result IDs
without requiring the pre-encrypted hash from the web interface.
2025-12-04 15:17:13 +03:00
saidsurucu 91ad04cf09 Fix: Catch all Playwright errors for curl fallback
ImportError only catches import failures. Browser launch errors
(executable not found) are runtime exceptions. Changed to catch
all Exception types to properly fallback to curl.
2025-12-04 14:44:39 +03:00
saidsurucu 5cec0df785 Add curl fallback for KİK document retrieval
Python SSL libraries (httpx, requests, urllib) fail with SSL handshake
errors against ekap.kik.gov.tr legacy server. curl uses different SSL
implementation (LibreSSL) that works.

Changes:
- Add subprocess + shutil imports
- Replace httpx fallback with curl fallback in get_document_markdown
- curl uses -k (insecure), -s (silent), -L (follow redirects) flags
- Enables KİK document retrieval on FastMCP Cloud without Playwright
2025-12-04 14:36:07 +03:00
saidsurucu d51f11c7ba Use setup.py post-install hook for Playwright Chromium installation
- Remove runtime subprocess install (FastMCP Cloud doesn't allow disk writes)
- Add setup.py with cmdclass hooks to install Chromium at build time
- Rename requirements.txt so FastMCP Cloud uses pyproject.toml instead
2025-12-04 13:34:10 +03:00
saidsurucu b207b16ef7 Auto-install Playwright Chromium at server startup for cloud deployments 2025-12-04 13:25:53 +03:00
saidsurucu f47147ba44 Add postinstall.sh for Playwright Chromium installation 2025-12-04 13:20:37 +03:00
86 changed files with 2503 additions and 18554 deletions
+9
View File
@@ -70,6 +70,15 @@ JWT_SECRET_KEY=your_jwt_secret_key_here
# MAX_REQUESTS_PER_MINUTE=60
# BURST_CAPACITY=20
# =============================================================================
# SEMANTIC SEARCH SETTINGS (Optional)
# =============================================================================
# OpenRouter API Key for semantic search functionality
# Get your API key from: https://openrouter.ai/keys
# If not set, semantic search tool will be disabled
OPENROUTER_API_KEY=sk-or-v1-your_openrouter_api_key_here
# =============================================================================
# USAGE INSTRUCTIONS
# =============================================================================
+2 -2
View File
@@ -1,5 +1,5 @@
# -------- BASE IMAGE (includes Chromium & deps) ----------------------------
FROM mcr.microsoft.com/playwright/python:v1.53.0-noble
# -------- BASE IMAGE ---------------------------------------------------------
FROM python:3.12-slim
# -------- Runtime setup ----------------------------------------------------
WORKDIR /app
+57 -2
View File
@@ -154,10 +154,65 @@ Yargı MCP'yi Gemini CLI ile kullanmak için:
</details>
---
<details>
<summary>🧠 <strong>Semantik Arama (Opsiyonel - OpenRouter API)</strong></summary>
Yargı MCP, **semantik arama** özelliği ile kararları anlamsal olarak sıralayabilir. Bu özellik opsiyoneldir ve `OPENROUTER_API_KEY` ayarlandığında otomatik olarak etkinleşir.
### Semantik Arama Nasıl Çalışır?
1. `initial_keyword` ile Bedesten API'den 100 karar çekilir
2. `query` ile bu kararlar embedding modeli kullanılarak anlamsal olarak sıralanır
3. En alakalı kararlar döndürülür
### OpenRouter API Anahtarı Alma
1. [OpenRouter](https://openrouter.ai/) sitesine gidin
2. Hesap oluşturun ve API anahtarı alın (ücretsiz kredi ile başlayabilirsiniz)
### Claude Desktop için Yapılandırma
```json
{
"mcpServers": {
"Yargı MCP": {
"command": "uvx",
"args": ["yargi-mcp"],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-xxx..."
}
}
}
}
```
### 5ire için Yapılandırma
Tool ayarlarında **Environment Variables** alanına ekleyin:
```
OPENROUTER_API_KEY=sk-or-v1-xxx...
```
### Gemini CLI için Yapılandırma
```json
{
"mcpServers": {
"yargi_mcp": {
"command": "uvx",
"args": ["yargi-mcp"],
"env": {
"OPENROUTER_API_KEY": "sk-or-v1-xxx..."
}
}
}
}
```
> 💡 **Not:** `OPENROUTER_API_KEY` ayarlanmazsa semantik arama aracı görünmez, diğer 19 araç normal şekilde çalışmaya devam eder.
</details>
<details>
<summary>🛠️ <strong>Kullanılabilir Araçlar (MCP Tools)</strong></summary>
Bu FastMCP sunucusu **19 optimize edilmiş MCP aracı** sunar (token verimliliği için optimize edilmiş):
Bu FastMCP sunucusu **19 temel MCP aracı** + **1 opsiyonel semantik arama aracı** sunar (token verimliliği için optimize edilmiş):
### **Yargıtay Araçları (Birleşik Bedesten API - Token Optimized)**
*Not: Yargıtay araçları token verimliliği için birleşik Bedesten API'ye entegre edilmiştir*
@@ -222,7 +277,7 @@ Bu FastMCP sunucusu **19 optimize edilmiş MCP aracı** sunar (token verimliliğ
**GENEL İSTATİSTİKLER:**
- **Toplam Mahkeme/Kurum:** 13 farklı hukuki kurum (KVKK dahil)
- **Toplam MCP Tool:** 19 optimize edilmiş arama ve belge getirme aracı
- **Toplam MCP Tool:** 19 temel araç + 1 opsiyonel semantik arama aracı
- **Daire/Kurul Filtreleme:** 87 farklı seçenek (52 Yargıtay + 27 Danıştay + 8 Sayıştay)
- **Tarih Filtreleme:** Birleşik Bedesten API aracında ISO 8601 formatında tam tarih aralığı desteği
- **Kesin Cümle Arama:** Birleşik Bedesten API aracında çift tırnak ile tam cümle arama (`"\"mülkiyet kararı\""` formatı)
File diff suppressed because it is too large Load Diff
+97 -74
View File
@@ -1,14 +1,21 @@
# kik_mcp_module/client_v2.py
import httpx
import requests
import logging
import uuid
import base64
import ssl
import os
from typing import Optional
from datetime import datetime
# Cryptography imports for AES-256-CBC encryption of document IDs
try:
from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes
from cryptography.hazmat.backends import default_backend
HAS_CRYPTOGRAPHY = True
except ImportError:
HAS_CRYPTOGRAPHY = False
from .models_v2 import (
KikV2DecisionType, KikV2SearchPayload, KikV2SearchPayloadDk, KikV2SearchPayloadMk,
KikV2RequestData, KikV2QueryRequest, KikV2KeyValuePair,
@@ -35,6 +42,54 @@ class KikV2ApiClient:
KikV2DecisionType.MAHKEME: "/b_ihalearaclari/api/KurulKararlari/GetKurulKararlariMk"
}
# AES-256-CBC encryption key for document ID encryption (reverse engineered from ekapv2.kik.gov.tr Angular app)
# This key is used to encrypt numeric gundemMaddesiId values to 64-character hex hashes for document URLs
DOCUMENT_ID_ENCRYPTION_KEY = bytes([
236, 193, 164, 43, 12, 135, 121, 170, 4, 244, 123, 219, 82, 158, 124, 174,
174, 228, 219, 174, 208, 104, 174, 120, 32, 76, 250, 4, 143, 159, 211, 176
])
@staticmethod
def encrypt_document_id(numeric_id: str) -> str:
"""
Encrypt a numeric KİK gundemMaddesiId to the 64-character hex hash
used in document URLs.
Algorithm: AES-256-CBC with PKCS7 padding
Output format: IV (16 bytes hex) + Ciphertext (16 bytes hex) = 64 chars
Args:
numeric_id: The numeric document ID from search results (e.g., "177280")
Returns:
64-character hex string for use in document URL KararId parameter
"""
if not HAS_CRYPTOGRAPHY:
raise ImportError("cryptography library required for document ID encryption")
# Generate random IV (16 bytes)
iv = os.urandom(16)
# Create AES-CBC cipher with the encryption key
cipher = Cipher(
algorithms.AES(KikV2ApiClient.DOCUMENT_ID_ENCRYPTION_KEY),
modes.CBC(iv),
backend=default_backend()
)
encryptor = cipher.encryptor()
# Encode plaintext and apply PKCS7 padding
plaintext = numeric_id.encode('utf-8')
block_size = 16
padding_len = block_size - (len(plaintext) % block_size)
padded_plaintext = plaintext + bytes([padding_len] * padding_len)
# Encrypt
ciphertext = encryptor.update(padded_plaintext) + encryptor.finalize()
# Return IV + ciphertext as lowercase hex (64 characters total)
return iv.hex() + ciphertext.hex()
def __init__(self, request_timeout: float = 60.0):
# Create SSL context with legacy server support
ssl_context = ssl.create_default_context()
@@ -270,7 +325,7 @@ class KikV2ApiClient:
This method uses a two-step process:
1. Call GetSorgulamaUrl endpoint to get the actual document URL
2. Use Playwright to navigate to that URL and extract content
2. Use httpx to fetch the document content
Args:
document_id: The gundemMaddesiId from search results
@@ -320,92 +375,60 @@ class KikV2ApiClient:
error_message="Could not get document URL from GetSorgulamaUrl API"
)
# Construct full document URL with the actual document ID
document_url = f"{base_document_url}?KararId={document_id}"
# If document_id is numeric, encrypt it to get the KararId hash
# The web interface uses AES-256-CBC encrypted hashes for document URLs
karar_id = document_id
if document_id.isdigit():
try:
karar_id = self.encrypt_document_id(document_id)
logger.info(f"KikV2ApiClient: Encrypted numeric ID {document_id} to hash: {karar_id}")
except Exception as enc_error:
logger.warning(f"KikV2ApiClient: Could not encrypt document ID, using as-is: {enc_error}")
# Construct full document URL with the encrypted KararId
document_url = f"{base_document_url}?KararId={karar_id}"
logger.info(f"KikV2ApiClient: Step 2 - Retrieved document URL: {document_url}")
except Exception as e:
logger.error(f"KikV2ApiClient: Error getting document URL for ID {document_id}: {str(e)}")
# Fallback to old method if GetSorgulamaUrl fails
document_url = f"https://ekap.kik.gov.tr/EKAP/Vatandas/KurulKararGoster.aspx?KararId={document_id}"
# Also encrypt numeric IDs in fallback path
karar_id = document_id
if document_id.isdigit():
try:
karar_id = self.encrypt_document_id(document_id)
logger.info(f"KikV2ApiClient: Encrypted numeric ID in fallback: {karar_id}")
except Exception as enc_error:
logger.warning(f"KikV2ApiClient: Could not encrypt in fallback: {enc_error}")
document_url = f"https://ekap.kik.gov.tr/EKAP/Vatandas/KurulKararGoster.aspx?KararId={karar_id}"
logger.info(f"KikV2ApiClient: Falling back to direct URL: {document_url}")
try:
# Step 2: Use Playwright to get the actual document content
logger.info(f"KikV2ApiClient: Step 2 - Using Playwright to retrieve document from: {document_url}")
# Step 2: Use httpx to get the document content
logger.info(f"KikV2ApiClient: Step 2 - Using httpx to retrieve document from: {document_url}")
try:
from playwright.async_api import async_playwright
# Create a separate httpx client for document retrieval with HTML headers
doc_ssl_context = ssl.create_default_context()
doc_ssl_context.check_hostname = False
doc_ssl_context.verify_mode = ssl.CERT_NONE
if hasattr(ssl, 'OP_LEGACY_SERVER_CONNECT'):
doc_ssl_context.options |= ssl.OP_LEGACY_SERVER_CONNECT
doc_ssl_context.set_ciphers('ALL:!aNULL:!eNULL:!EXPORT:!DES:!RC4:!MD5:!PSK:!SRP:!CAMELLIA')
async with async_playwright() as p:
# Launch browser
browser = await p.chromium.launch(
headless=True,
args=['--no-sandbox', '--disable-dev-shm-usage']
)
page = await browser.new_page(
user_agent="Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/139.0.0.0 Safari/537.36"
)
# Navigate to document page with longer timeout for JS loading
await page.goto(document_url, wait_until="networkidle", timeout=15000)
# Wait for the document content to load (KİK pages might need more time for JS execution)
await page.wait_for_timeout(3000)
# Wait for Angular/Zone.js to finish loading and document to be ready
try:
# Wait for Angular zone to be available (this JavaScript code you showed)
await page.wait_for_function(
"typeof Zone !== 'undefined' && Zone.current",
timeout=10000
)
# Wait for network to be idle after Angular bootstrap
await page.wait_for_load_state("networkidle", timeout=10000)
# Wait for specific document content to appear
await page.wait_for_function(
"""
document.body.textContent.length > 5000 &&
(document.body.textContent.includes('Karar') ||
document.body.textContent.includes('KURUL') ||
document.body.textContent.includes('Gündem') ||
document.body.textContent.includes('Toplantı'))
""",
timeout=15000
)
logger.info("KikV2ApiClient: Angular document content loaded successfully")
except Exception as e:
logger.warning(f"KikV2ApiClient: Angular content loading timed out, proceeding anyway: {str(e)}")
# Give a bit more time for any remaining content to load
await page.wait_for_timeout(5000)
# Get page content
html_content = await page.content()
await browser.close()
logger.info(f"KikV2ApiClient: Retrieved content via Playwright, length: {len(html_content)}")
except ImportError:
logger.info("KikV2ApiClient: Playwright not available, falling back to httpx")
# Fallback to httpx
response = await self.http_client.get(
document_url,
async with httpx.AsyncClient(
verify=doc_ssl_context,
headers={
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "tr,en-US;q=0.5",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/139.0.0.0 Safari/537.36",
"Referer": "https://ekap.kik.gov.tr/",
"Cache-Control": "no-cache"
}
)
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/139.0.0.0 Safari/537.36"
},
timeout=60.0,
follow_redirects=True
) as doc_client:
response = await doc_client.get(document_url)
response.raise_for_status()
html_content = response.text
logger.info(f"KikV2ApiClient: Retrieved content via httpx, length: {len(html_content)}")
# Convert HTML to Markdown using MarkItDown with BytesIO
try:
-74
View File
@@ -1,74 +0,0 @@
# kik_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl, computed_field, ConfigDict
from typing import List, Optional
from enum import Enum
import base64 # Base64 encoding/decoding için
class KikKararTipi(str, Enum):
"""Enum for KIK (Public Procurement Authority) Decision Types."""
UYUSMAZLIK = "rbUyusmazlik"
DUZENLEYICI = "rbDuzenleyici"
MAHKEME = "rbMahkeme"
class KikSearchRequest(BaseModel):
"""Model for KIK Decision search criteria."""
karar_tipi: KikKararTipi = Field(KikKararTipi.UYUSMAZLIK, description="Type")
karar_no: str = Field("", description="No")
karar_tarihi_baslangic: str = Field("", description="Start", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
karar_tarihi_bitis: str = Field("", description="End", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
resmi_gazete_sayisi: str = Field("", description="Gazette")
resmi_gazete_tarihi: str = Field("", description="Date", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
basvuru_konusu_ihale: str = Field("", description="Subject")
basvuru_sahibi: str = Field("", description="Applicant")
ihaleyi_yapan_idare: str = Field("", description="Entity")
yil: str = Field("", description="Year")
karar_metni: str = Field("", description="Text")
page: int = Field(1, ge=1, description="Page")
class KikDecisionEntry(BaseModel):
"""Represents a single decision entry from KIK search results."""
preview_event_target: str = Field(..., description="Event target")
karar_no_str: str = Field(..., alias="kararNo", description="Decision number")
karar_tipi: KikKararTipi = Field(..., description="Decision type")
karar_tarihi_str: str = Field(..., alias="kararTarihi", description="Date")
idare_str: str = Field("", alias="idare", description="Entity")
basvuru_sahibi_str: str = Field("", alias="basvuruSahibi", description="Applicant")
ihale_konusu_str: str = Field("", alias="ihaleKonusu", description="Subject")
@computed_field
@property
def karar_id(self) -> str:
"""
A Base64 encoded unique ID for the decision, combining decision type and number.
Format before encoding: "{karar_tipi.value}|{karar_no_str}"
"""
combined_key = f"{self.karar_tipi.value}|{self.karar_no_str}"
return base64.b64encode(combined_key.encode('utf-8')).decode('utf-8')
model_config = ConfigDict(populate_by_name=True)
class KikSearchResult(BaseModel):
"""Model for KIK search results."""
decisions: List[KikDecisionEntry]
total_records: int = 0
current_page: int = 1
class KikDocumentMarkdown(BaseModel):
"""
KIK decision document, with Markdown content potentially paginated.
"""
retrieved_with_karar_id: Optional[str] = Field(None, description="Request ID")
retrieved_karar_no: Optional[str] = Field(None, description="Decision number")
retrieved_karar_tipi: Optional[KikKararTipi] = Field(None, description="Decision type")
karar_id_param_from_url: Optional[str] = Field(None, alias="kararIdParam", description="Internal ID")
markdown_chunk: Optional[str] = Field(None, description="Content")
source_url: Optional[str] = Field(None, description="Source URL")
error_message: Optional[str] = Field(None, description="Error")
current_page: int = Field(1, description="Page")
total_pages: int = Field(1, description="Total pages")
is_paginated: bool = Field(False, description="Paginated")
full_content_char_count: Optional[int] = Field(None, description="Char count")
model_config = ConfigDict(populate_by_name=True)
+302 -107
View File
@@ -2,14 +2,12 @@
import asyncio
import atexit
import logging
import os
import httpx
import json
import time
from collections import defaultdict
from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, List, Literal, Any, Union
import urllib.parse
from pydantic import HttpUrl, Field
from typing import Optional, Dict, List, Literal, Any
from fastmcp.server.middleware import Middleware, MiddlewareContext
# Optional tiktoken import for token counting
@@ -150,7 +148,7 @@ class TokenCountingMiddleware(Middleware):
return result
except Exception as e:
except Exception:
duration_ms = (time.perf_counter() - start_time) * 1000
self.log_token_usage("tool_call_error", input_tokens, 0,
tool_name, duration_ms)
@@ -180,7 +178,7 @@ class TokenCountingMiddleware(Middleware):
return result
except Exception as e:
except Exception:
duration_ms = (time.perf_counter() - start_time) * 1000
self.log_token_usage("resource_read_error", 0, 0,
resource_uri, duration_ms)
@@ -210,7 +208,7 @@ class TokenCountingMiddleware(Middleware):
return result
except Exception as e:
except Exception:
duration_ms = (time.perf_counter() - start_time) * 1000
self.log_token_usage("prompt_get_error", 0, 0,
prompt_name, duration_ms)
@@ -251,71 +249,56 @@ def create_app(auth=None):
# --- Module Imports ---
from yargitay_mcp_module.client import YargitayOfficialApiClient
from yargitay_mcp_module.models import (
YargitayDetailedSearchRequest, YargitayDocumentMarkdown, CompactYargitaySearchResult,
YargitayBirimEnum, CleanYargitayDecisionEntry
)
from bedesten_mcp_module.client import BedestenApiClient
from bedesten_mcp_module.models import (
BedestenSearchRequest, BedestenSearchData,
BedestenDocumentMarkdown, BedestenCourtTypeEnum
)
from bedesten_mcp_module.enums import BirimAdiEnum
# Semantic Search Module Imports (conditional based on OPENROUTER_API_KEY)
from semantic_search.embedder import is_openrouter_available
SEMANTIC_SEARCH_AVAILABLE = is_openrouter_available()
if SEMANTIC_SEARCH_AVAILABLE:
from semantic_search.embedder import OpenRouterEmbedder
from semantic_search.vector_store import VectorStore
from semantic_search.processor import DocumentProcessor
logger.info("Semantic search enabled (OPENROUTER_API_KEY found)")
else:
logger.info("Semantic search disabled (OPENROUTER_API_KEY not set)")
from danistay_mcp_module.client import DanistayApiClient
from danistay_mcp_module.models import (
DanistayKeywordSearchRequest, DanistayDetailedSearchRequest,
DanistayDocumentMarkdown, CompactDanistaySearchResult
)
from emsal_mcp_module.client import EmsalApiClient
from emsal_mcp_module.models import (
EmsalSearchRequest, EmsalDocumentMarkdown, CompactEmsalSearchResult
EmsalSearchRequest, CompactEmsalSearchResult
)
from uyusmazlik_mcp_module.client import UyusmazlikApiClient
from uyusmazlik_mcp_module.models import (
UyusmazlikSearchRequest, UyusmazlikSearchResponse, UyusmazlikDocumentMarkdown,
UyusmazlikBolumEnum, UyusmazlikTuruEnum, UyusmazlikKararSonucuEnum
UyusmazlikSearchRequest, UyusmazlikBolumEnum, UyusmazlikTuruEnum, UyusmazlikKararSonucuEnum
)
from anayasa_mcp_module.client import AnayasaMahkemesiApiClient
from anayasa_mcp_module.bireysel_client import AnayasaBireyselBasvuruApiClient
from anayasa_mcp_module.unified_client import AnayasaUnifiedClient
from anayasa_mcp_module.models import (
AnayasaNormDenetimiSearchRequest,
AnayasaSearchResult,
AnayasaDocumentMarkdown,
AnayasaBireyselReportSearchRequest,
AnayasaBireyselReportSearchResult,
AnayasaBireyselBasvuruDocumentMarkdown,
AnayasaUnifiedSearchRequest,
AnayasaUnifiedSearchResult,
AnayasaUnifiedDocumentMarkdown,
# Removed enum imports - now using Literal strings in models
)
# KIK v2 Module Imports (New API)
from kik_mcp_module.client_v2 import KikV2ApiClient
from kik_mcp_module.models_v2 import KikV2DecisionType
from kik_mcp_module.models_v2 import (
KikV2SearchResult,
KikV2DocumentMarkdown
)
from rekabet_mcp_module.client import RekabetKurumuApiClient
from rekabet_mcp_module.models import (
RekabetKurumuSearchRequest,
RekabetSearchResult,
RekabetDocument,
RekabetKararTuruGuidEnum
)
from sayistay_mcp_module.client import SayistayApiClient
from sayistay_mcp_module.models import (
GenelKurulSearchRequest, GenelKurulSearchResponse,
TemyizKuruluSearchRequest, TemyizKuruluSearchResponse,
DaireSearchRequest, DaireSearchResponse,
SayistayDocumentMarkdown,
SayistayUnifiedSearchRequest, SayistayUnifiedSearchResult,
SayistayUnifiedDocumentMarkdown
SayistayUnifiedSearchRequest
)
from sayistay_mcp_module.enums import DaireEnum, KamuIdaresiTuruEnum, WebKararKonusuEnum
from sayistay_mcp_module.unified_client import SayistayUnifiedClient
# KVKK Module Imports
@@ -329,14 +312,11 @@ from kvkk_mcp_module.models import (
# BDDK Module Imports
from bddk_mcp_module.client import BddkApiClient
from bddk_mcp_module.models import (
BddkSearchRequest,
BddkSearchResult,
BddkDocumentMarkdown
BddkSearchRequest
)
# Create a placeholder app that will be properly initialized after tools are defined
from fastmcp import FastMCP
# MCP app for Turkish legal databases with explicit capabilities
app = FastMCP(
@@ -646,7 +626,7 @@ async def search_emsal_detailed_decisions(
page_size=page_size
)
logger.info(f"Tool 'search_emsal_detailed_decisions' called.")
logger.info("Tool 'search_emsal_detailed_decisions' called.")
try:
api_response = await emsal_client_instance.search_detailed_decisions(search_query)
if api_response.data:
@@ -658,8 +638,8 @@ async def search_emsal_detailed_decisions(
).model_dump()
logger.warning("API response for Emsal search did not contain expected data structure.")
return CompactEmsalSearchResult(decisions=[], total_records=0, requested_page=search_query.page_number, page_size=search_query.page_size).model_dump()
except Exception as e:
logger.exception(f"Error in tool 'search_emsal_detailed_decisions'.")
except Exception:
logger.exception("Error in tool 'search_emsal_detailed_decisions'.")
raise
@app.tool(
@@ -676,8 +656,8 @@ async def get_emsal_document_markdown(id: str) -> Dict[str, Any]:
try:
result = await emsal_client_instance.get_decision_document_as_markdown(id)
return result.model_dump()
except Exception as e:
logger.exception(f"Error in tool 'get_emsal_document_markdown'.")
except Exception:
logger.exception("Error in tool 'get_emsal_document_markdown'.")
raise
# --- MCP Tools for Uyusmazlik ---
@@ -745,12 +725,12 @@ async def search_uyusmazlik_decisions(
not_hepsi=not_hepsi
)
logger.info(f"Tool 'search_uyusmazlik_decisions' called.")
logger.info("Tool 'search_uyusmazlik_decisions' called.")
try:
result = await uyusmazlik_client_instance.search_decisions(search_params)
return result.model_dump()
except Exception as e:
logger.exception(f"Error in tool 'search_uyusmazlik_decisions'.")
except Exception:
logger.exception("Error in tool 'search_uyusmazlik_decisions'.")
raise
@app.tool(
@@ -770,8 +750,8 @@ async def get_uyusmazlik_document_markdown_from_url(
try:
result = await uyusmazlik_client_instance.get_decision_document_as_markdown(str(document_url))
return result.model_dump()
except Exception as e:
logger.exception(f"Error in tool 'get_uyusmazlik_document_markdown_from_url'.")
except Exception:
logger.exception("Error in tool 'get_uyusmazlik_document_markdown_from_url'.")
raise
# --- DEACTIVATED: MCP Tools for Anayasa Mahkemesi (Individual Tools) ---
@@ -862,8 +842,8 @@ async def search_anayasa_unified(
result = await anayasa_unified_client_instance.search_unified(request)
return json.dumps(result.model_dump(), ensure_ascii=False, indent=2)
except Exception as e:
logger.exception(f"Error in tool 'search_anayasa_unified'.")
except Exception:
logger.exception("Error in tool 'search_anayasa_unified'.")
raise
@app.tool(
@@ -884,8 +864,8 @@ async def get_anayasa_document_unified(
result = await anayasa_unified_client_instance.get_document_unified(document_url, page_number)
return json.dumps(result.model_dump(mode='json'), ensure_ascii=False, indent=2)
except Exception as e:
logger.exception(f"Error in tool 'get_anayasa_document_unified'.")
except Exception:
logger.exception("Error in tool 'get_anayasa_document_unified'.")
raise
# --- MCP Tools for KIK v2 (Kamu İhale Kurulu - New API) ---
@@ -972,23 +952,23 @@ async def search_kik_v2_decisions(
}
)
async def get_kik_v2_document_markdown(
document_id: str = Field(..., description="Document ID (gundemMaddesiId from search results)")
gundemMaddesiId: str = Field(..., description="gundemMaddesiId from search_kik_v2_decisions results")
) -> dict:
"""Get KİK decision document using the v2 API."""
"""Get KİK decision document in Markdown format."""
logger.info(f"Tool 'get_kik_v2_document_markdown' called for document ID: {document_id}")
logger.info(f"Tool 'get_kik_v2_document_markdown' called for gundemMaddesiId: {gundemMaddesiId}")
if not document_id or not document_id.strip():
if not gundemMaddesiId or not gundemMaddesiId.strip():
return {
"document_id": document_id,
"document_id": gundemMaddesiId,
"kararNo": "",
"markdown_content": "",
"source_url": "",
"error_message": "Document ID is required and must be a non-empty string"
"error_message": "gundemMaddesiId is required and must be a non-empty string"
}
try:
api_response = await kik_v2_client_instance.get_document_markdown(document_id)
api_response = await kik_v2_client_instance.get_document_markdown(gundemMaddesiId)
return {
"document_id": api_response.document_id,
@@ -999,9 +979,9 @@ async def get_kik_v2_document_markdown(
}
except Exception as e:
logger.exception(f"Error in KİK v2 document retrieval tool for document ID: {document_id}")
logger.exception(f"Error in KİK v2 document retrieval tool for gundemMaddesiId: {gundemMaddesiId}")
return {
"document_id": document_id,
"document_id": gundemMaddesiId,
"kararNo": "",
"markdown_content": "",
"source_url": "",
@@ -1060,7 +1040,7 @@ async def search_rekabet_kurumu_decisions(
result = await rekabet_client_instance.search_decisions(search_query)
return result.model_dump()
except Exception as e:
except Exception:
logger.exception("Error in tool 'search_rekabet_kurumu_decisions'.")
return RekabetSearchResult(decisions=[], retrieved_page_number=page, total_records_found=0, total_pages=0).model_dump()
@@ -1083,7 +1063,7 @@ async def get_rekabet_kurumu_document(
try:
result = await rekabet_client_instance.get_decision_document(karar_id, page_number=current_page_to_fetch)
return result.model_dump()
except Exception as e:
except Exception:
logger.exception(f"Error in tool 'get_rekabet_kurumu_document'. Karar ID: {karar_id}")
raise
@@ -1194,7 +1174,7 @@ For best results, use exact phrases with quotes for legal terms."""),
"page_size": pageSize,
"searched_courts": court_types
}
except Exception as e:
except Exception:
logger.exception("Error in tool 'search_bedesten_unified'")
raise
@@ -1216,10 +1196,258 @@ async def get_bedesten_document_markdown(
try:
return await bedesten_client_instance.get_document_as_markdown(documentId)
except Exception as e:
except Exception:
logger.exception("Error in tool 'get_kyb_bedesten_document_markdown'")
raise
# --- Semantic Search Tool (Conditional - requires OPENROUTER_API_KEY) ---
if SEMANTIC_SEARCH_AVAILABLE:
@app.tool(
description="Perform semantic search on Turkish legal decisions using OpenRouter Gemini embeddings for intelligent re-ranking",
annotations={
"readOnlyHint": True,
"openWorldHint": True,
"idempotentHint": True
}
)
async def search_bedesten_semantic(
initial_keyword: str = Field(..., description="""Bedesten API'den ilk sonuçları çekmek için anahtar kelime veya arama ifadesi.
Bu terim ile API'den 100 karar çekilir, sonra semantik sıralama yapılır.
ARAMA OPERATÖRLERİ:
• Basit arama: "muvazaa" (kelimeyi içeren kararlar)
• Tam eşleşme: "\"muris muvazaası\"" (tırnak içi aynen aranır)
• AND: "muvazaa AND tapu" (her iki terim zorunlu)
• OR: "ecrimisil OR kira" (en az biri yeterli)
• NOT: "muvazaa NOT miras" (muvazaa içeren ama miras içermeyen)
• Zorunlu: "+muvazaa tapu" (muvazaa zorunlu, tapu opsiyonel)
• Hariç: "muvazaa -miras" (muvazaa içeren, miras hariç)
ÖRNEKLER:
"muvazaa" - geniş arama
"\"muris muvazaası\"" - tam ifade
"muvazaa AND tapu AND iptal" - tüm terimler zorunlu
"ecrimisil OR haksız işgal" - alternatifli arama"""),
query: str = Field(..., description="""Semantik benzerlik için DETAYLI arama sorgusu.
initial_keyword ile bulunan kararlar bu sorguya göre anlamsal olarak sıralanır.
ÖNEMLİ: Embedding modeli anlamlı cümleler bekler, anahtar kelimeler DEĞİL.
Aradığınız hukuki meseleyi CÜMLE olarak yazın.
DOĞRU KULLANIM:
"Mirasçının muvazaalı satış işlemine karşı tapu iptali ve tescil davası açması"
"Taşınmazın fiili kullanımı ve zilyetlik durumunun değerlendirilmesi"
"İş sözleşmesinin feshinde kıdem tazminatı hesaplama yöntemi"
YANLIŞ KULLANIM:
"muvazaa tapu iptal" (sadece kelimeler, cümle değil)
"kıdem tazminat hesap" (bağlamsız kelimeler)
İPUCU: Ne arıyorsanız onu bir cümle olarak ifade edin."""),
court_types: List[BedestenCourtTypeEnum] = Field(
default=["YARGITAYKARARI", "DANISTAYKARAR", "YERELHUKUK", "ISTINAFHUKUK", "KYB"],
description="Court types to search: YARGITAYKARARI, DANISTAYKARAR, YERELHUKUK, ISTINAFHUKUK, KYB (default: all)"
),
top_k: int = Field(10, ge=1, le=50, description="Number of top results to return (1-50)")
) -> Dict[str, Any]:
"""
Perform semantic search on Turkish legal decisions using OpenRouter API.
This tool:
1. Searches Bedesten API with initial keyword (retrieves 100 results)
2. Fetches full document content for each result
3. Generates embeddings using Google's Gemini Embedding model via OpenRouter
4. Performs semantic similarity search with the query
5. Returns re-ranked results based on semantic relevance
Benefits over keyword search:
- Better understanding of context and meaning
- Finds semantically similar documents even with different wording
- More accurate ranking based on relevance
- Supports multilingual queries (100+ languages)
Note: Requires OPENROUTER_API_KEY environment variable to be set.
"""
logger.info(f"Semantic search tool called with initial_keyword: {initial_keyword}, query: {query}")
try:
# Initialize components
embedder = OpenRouterEmbedder()
vector_store = VectorStore(dimension=3072) # Gemini embedding dimension
processor = DocumentProcessor(chunk_size=1500, chunk_overlap=300)
# Step 1: Initial keyword search to get document IDs
logger.info(f"Step 1: Searching Bedesten API with keyword: {initial_keyword}")
all_decisions = []
# Search each court type
for court_type in court_types:
try:
per_court_limit = max(20, 100 // len(court_types))
search_results = await bedesten_client_instance.search_documents(
BedestenSearchRequest(
data=BedestenSearchData(
phrase=initial_keyword,
itemTypeList=[court_type],
pageSize=per_court_limit,
pageNumber=1
)
)
)
if search_results.data and search_results.data.emsalKararList:
all_decisions.extend(search_results.data.emsalKararList)
logger.info(f"Found {len(search_results.data.emsalKararList)} results from {court_type}")
except Exception as e:
logger.warning(f"Error searching {court_type}: {e}")
if not all_decisions:
logger.warning("No documents found from initial search")
return {
"status": "no_results",
"message": "No documents found matching the initial keyword",
"results": []
}
logger.info(f"Total documents found: {len(all_decisions)}")
# Step 2: Fetch document content and process
logger.info("Step 2: Fetching and processing document content...")
documents_data = []
failed_fetches = 0
decisions_to_process = all_decisions[:100]
for i, decision in enumerate(decisions_to_process):
try:
doc = await bedesten_client_instance.get_document_as_markdown(decision.documentId)
if doc.markdown_content:
metadata = {
"document_id": decision.documentId,
"birim_adi": decision.birimAdi,
"esas_no": decision.esasNo,
"karar_no": decision.kararNo,
"karar_tarihi": decision.kararTarihiStr,
"court_type": decision.itemType.name if decision.itemType else None
}
chunks = processor.process_document(
document_id=decision.documentId,
text=doc.markdown_content,
metadata=metadata
)
if chunks:
full_text = " ".join([chunk.text for chunk in chunks])
documents_data.append({
"id": decision.documentId,
"text": full_text[:3000],
"metadata": metadata
})
if (i + 1) % 10 == 0:
logger.info(f"Processed {i + 1}/{len(decisions_to_process)} documents")
except Exception as e:
logger.warning(f"Failed to fetch document {decision.documentId}: {e}")
failed_fetches += 1
if not documents_data:
logger.warning("No documents could be processed")
return {
"status": "processing_error",
"message": "Could not process any documents",
"results": []
}
logger.info(f"Successfully processed {len(documents_data)} documents, {failed_fetches} failed")
# Step 3: Generate embeddings
logger.info("Step 3: Generating embeddings...")
query_embedding = embedder.encode_query(query, task="search result")
doc_texts = [doc["text"] for doc in documents_data]
doc_titles = [doc["metadata"].get("birim_adi", "none") for doc in documents_data]
doc_embeddings = embedder.encode_documents(doc_texts, titles=doc_titles)
# No dimension reduction - using full 3072 dimensions
# Step 4: Add to vector store and search
logger.info("Step 4: Performing semantic search...")
doc_ids = [doc["id"] for doc in documents_data]
doc_metadatas = [doc["metadata"] for doc in documents_data]
vector_store.add_documents(
ids=doc_ids,
texts=doc_texts,
embeddings=doc_embeddings,
metadata=doc_metadatas
)
search_results = vector_store.search(
query_embedding=query_embedding,
top_k=top_k,
threshold=0.3
)
# Step 5: Format results
logger.info(f"Step 5: Formatting {len(search_results)} results")
formatted_results = []
for doc, score in search_results:
title_parts = []
if doc.metadata.get("birim_adi"):
title_parts.append(doc.metadata["birim_adi"])
if doc.metadata.get("esas_no"):
title_parts.append(f"Esas: {doc.metadata['esas_no']}")
if doc.metadata.get("karar_no"):
title_parts.append(f"Karar: {doc.metadata['karar_no']}")
if doc.metadata.get("karar_tarihi"):
title_parts.append(f"Tarih: {doc.metadata['karar_tarihi']}")
title = " - ".join(title_parts) if title_parts else f"Document {doc.id}"
formatted_results.append({
"document_id": doc.id,
"title": title,
"similarity_score": float(score),
"preview": doc.text[:500] + "..." if len(doc.text) > 500 else doc.text,
"metadata": doc.metadata,
"source_url": f"https://mevzuat.adalet.gov.tr/ictihat/{doc.id}"
})
stats = vector_store.get_stats()
return {
"status": "success",
"query": query,
"initial_keyword": initial_keyword,
"total_documents_processed": len(documents_data),
"embedding_dimension": 3072,
"results": formatted_results,
"stats": {
"documents_in_store": stats["num_documents"],
"memory_usage_mb": round(stats["memory_usage_mb"], 2),
"failed_fetches": failed_fetches
}
}
except Exception as e:
logger.exception(f"Error in semantic search: {e}")
return {
"status": "error",
"message": str(e),
"results": []
}
# --- MCP Tools for Sayıştay (Turkish Court of Accounts) ---
# DEACTIVATED TOOL - Use search_sayistay_unified instead
@@ -1406,7 +1634,7 @@ async def search_sayistay_unified(
)
result = await sayistay_unified_client_instance.search_unified(search_request)
return result.model_dump()
except Exception as e:
except Exception:
logger.exception("Error in tool 'search_sayistay_unified'")
raise
@@ -1431,7 +1659,7 @@ async def get_sayistay_document_unified(
try:
result = await sayistay_unified_client_instance.get_document_unified(decision_id, decision_type)
return result.model_dump()
except Exception as e:
except Exception:
logger.exception("Error in tool 'get_sayistay_document_unified'")
raise
@@ -2046,7 +2274,7 @@ async def search(
]
}
except Exception as e:
except Exception:
logger.exception("Error in ChatGPT Deep Research search tool")
# Return partial results if any were found
if results:
@@ -2185,42 +2413,12 @@ async def fetch(
doc = await bedesten_client_instance.get_document_as_markdown(doc_id)
"""
except Exception as e:
except Exception:
logger.exception(f"Error fetching ChatGPT Deep Research document {id}")
raise
# --- Token Metrics Tool Removed for Optimization ---
def ensure_playwright_browsers():
"""Ensure Playwright browsers are installed for KIK tool functionality."""
try:
import subprocess
import os
# Check if chromium is already installed
chromium_path = os.path.expanduser("~/Library/Caches/ms-playwright/chromium-1179")
if os.path.exists(chromium_path):
logger.info("Playwright Chromium browser already installed.")
return
logger.info("Installing Playwright Chromium browser for KIK tool...")
result = subprocess.run(
["python", "-m", "playwright", "install", "chromium"],
capture_output=True,
text=True,
timeout=300 # 5 minutes timeout
)
if result.returncode == 0:
logger.info("Playwright Chromium browser installed successfully.")
else:
logger.warning(f"Failed to install Playwright browser: {result.stderr}")
logger.warning("KIK tool may not work properly without Playwright browsers.")
except Exception as e:
logger.warning(f"Could not auto-install Playwright browsers: {e}")
logger.warning("KIK tool may not work properly. Manual installation: 'playwright install chromium'")
def main():
# Initialize the app properly with create_app()
global app
@@ -2229,14 +2427,11 @@ def main():
logger.info(f"Starting {app.name} server via main() function...")
# logger.info(f"Logs will be written to: {LOG_FILE_PATH}") # File logging disabled
# Ensure Playwright browsers are installed
ensure_playwright_browsers()
try:
app.run()
except KeyboardInterrupt:
logger.info("Server shut down by user (KeyboardInterrupt).")
except Exception as e:
except Exception:
logger.exception("Server failed to start or crashed.")
finally:
logger.info(f"{app.name} server has shut down.")
+5 -3
View File
@@ -1,6 +1,6 @@
[project]
name = "yargi-mcp"
version = "0.1.9"
version = "0.2.0"
description = "MCP Server For Turkish Legal Databases"
readme = "README.md"
requires-python = ">=3.11"
@@ -25,10 +25,12 @@ dependencies = [
"markitdown[pdf]>=0.1.1",
"pydantic>=2.11.4",
"aiohttp>=3.11.18",
"playwright>=1.52.0",
"fastmcp>=2.10.5",
"pypdf>=5.5.0",
"fastapi>=0.115.14",
"cryptography>=44.0.0",
"openai>=1.0.0",
"numpy>=1.24.0",
]
[project.optional-dependencies]
@@ -59,7 +61,7 @@ yargi-mcp = "mcp_server_main:main"
py-modules = ["mcp_server_main", "mcp_auth_factory", "mcp_auth_http_adapter", "asgi_app", "fastapi_app", "starlette_app", "run_asgi", "stripe_webhook"]
[tool.setuptools.packages.find]
include = ["*_mcp_module", "mcp_auth"]
include = ["*_mcp_module", "mcp_auth", "semantic_search"]
[build-system]
requires = ["setuptools>=65.0", "wheel"]
-11
View File
@@ -1,11 +0,0 @@
fastmcp
httpx
beautifulsoup4
markitdown[pdf]
pydantic
aiohttp
playwright
pypdf
fastapi>=0.115.14
uvicorn[standard]>=0.30.0
starlette>=0.37.0
-186
View File
@@ -1,186 +0,0 @@
# flyctl launch added from .gitignore
# Byte-compiled / optimized / DLL files
**/__pycache__
**/*.py[cod]
**/*$py.class
# C extensions
**/*.so
# Distribution / packaging
**/.Python
**/build
**/develop-eggs
**/dist
**/downloads
**/eggs
**/.eggs
**/lib
**/lib64
**/parts
**/sdist
**/var
**/wheels
**/share/python-wheels
**/*.egg-info
**/.installed.cfg
**/*.egg
**/MANIFEST
# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
**/*.manifest
**/*.spec
# Installer logs
**/pip-log.txt
**/pip-delete-this-directory.txt
# Unit test / coverage reports
**/htmlcov
**/.tox
**/.nox
**/.coverage
**/.coverage.*
**/.cache
**/nosetests.xml
**/coverage.xml
**/*.cover
**/*.py,cover
**/.hypothesis
**/.pytest_cache
**/cover
# Translations
**/*.mo
**/*.pot
# Django stuff:
**/*.log
**/local_settings.py
**/db.sqlite3
**/db.sqlite3-journal
# Flask stuff:
**/instance
**/.webassets-cache
# Scrapy stuff:
**/.scrapy
# Sphinx documentation
**/docs/_build
# PyBuilder
**/.pybuilder
**/target
# Jupyter Notebook
**/.ipynb_checkpoints
# IPython
**/profile_default
**/ipython_config.py
# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version
# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock
# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock
# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
#pdm.lock
# pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
# in version control.
# https://pdm.fming.dev/#use-with-ide
**/.pdm.toml
# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
**/__pypackages__
# Celery stuff
**/celerybeat-schedule
**/celerybeat.pid
# SageMath parsed files
**/*.sage.py
# Environments
**/.env
**/.venv
**/env
**/venv
**/ENV
**/env.bak
**/venv.bak
# Spyder project settings
**/.spyderproject
**/.spyproject
# Rope project settings
**/.ropeproject
# mkdocs documentation
site
# mypy
**/.mypy_cache
**/.dmypy.json
**/dmypy.json
# Pyre type checker
**/.pyre
# pytype static type analyzer
**/.pytype
# Cython debug symbols
**/cython_debug
# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/
**/.DS_Store
**/hello.py
**/*.html
**/fast-mcp-docs.md
# Debug and test files
**/debug_*
**/test_*
**/CLAUDE.md
# ASGI/Deployment files
**/ssl
**/*.pem
**/*.key
**/*.crt
# Docker volumes
**/redis-data
# Production logs
**/logs/*.log.*
**/Dockerfile
**/Dockerfile
fly.toml
-103
View File
@@ -1,103 +0,0 @@
# OAuth Configuration for Clerk + Google
# Copy this file to .env and fill in your actual values
# =============================================================================
# AUTHENTICATION SETTINGS
# =============================================================================
# Enable/disable authentication (set to "true" to enable OAuth)
ENABLE_AUTH=false
# =============================================================================
# CLERK CONFIGURATION
# =============================================================================
# Clerk API keys (get from https://dashboard.clerk.com/)
CLERK_SECRET_KEY=sk_test_your_secret_key_here
CLERK_PUBLISHABLE_KEY=pk_test_your_publishable_key_here
# OAuth Redirect URLs
CLERK_OAUTH_REDIRECT_URL=http://localhost:8000/auth/callback
CLERK_FRONTEND_URL=http://localhost:3000
# Clerk domain issuer (usually auto-configured)
CLERK_ISSUER=https://your-clerk-domain.clerk.accounts.dev
CLERK_DOMAIN=your-clerk-domain
# =============================================================================
# GOOGLE OAUTH SETTINGS
# =============================================================================
# Note: Google OAuth is configured through Clerk dashboard
# You need to:
# 1. Go to Clerk Dashboard > Social Connections
# 2. Enable Google provider
# 3. Add your Google OAuth client ID and secret
# 4. Configure redirect URIs in Google Console
# =============================================================================
# STRIPE CONFIGURATION (for payments/subscriptions)
# =============================================================================
STRIPE_SECRET=sk_test_your_stripe_secret_key_here
STRIPE_WEBHOOK_SECRET=whsec_your_webhook_secret_here
# =============================================================================
# SERVER CONFIGURATION
# =============================================================================
# CORS origins (comma-separated list)
ALLOWED_ORIGINS=http://localhost:3000,http://localhost:8000,https://yourdomain.com
# Server settings
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
# Base URL for the application (used for OAuth callbacks and API URLs)
BASE_URL=http://localhost:8000
# JWT Secret for MCP token generation
JWT_SECRET_KEY=your_jwt_secret_key_here
# =============================================================================
# MCP SERVER SETTINGS
# =============================================================================
# Additional MCP server configuration can go here
# For example, rate limiting, feature flags, etc.
# Example: Rate limiting
# MAX_REQUESTS_PER_MINUTE=60
# BURST_CAPACITY=20
# =============================================================================
# USAGE INSTRUCTIONS
# =============================================================================
# 1. Copy this file to .env:
# cp .env.example .env
# 2. Get Clerk credentials:
# - Sign up at https://clerk.com/
# - Create a new application
# - Go to API Keys tab
# - Copy Secret Key and Publishable Key
# 3. Configure Google OAuth in Clerk:
# - In Clerk Dashboard, go to Social Connections
# - Enable Google provider
# - Get Google OAuth credentials from Google Console
# - Add redirect URI: http://localhost:8000/auth/callback
# 4. Update OAuth URLs:
# - Set CLERK_OAUTH_REDIRECT_URL to your callback URL
# - Set CLERK_FRONTEND_URL to your frontend application URL
# 5. Enable authentication:
# - Set ENABLE_AUTH=true
# 6. Test the OAuth flow:
# - Start server: uvicorn asgi_app:app --reload
# - Visit: http://localhost:8000/auth/login
# - Complete OAuth flow with Google
# - Check: http://localhost:8000/auth/user
@@ -1,2 +0,0 @@
# Auto detect text files and perform LF normalization
* text=auto
@@ -1,37 +0,0 @@
name: Publish to PyPI
on:
release:
types: [published]
workflow_dispatch: # Manual trigger for testing
jobs:
pypi-publish:
name: Upload release to PyPI
runs-on: ubuntu-latest
environment:
name: pypi
url: https://pypi.org/p/yargi-mcp
permissions:
id-token: write # IMPORTANT: this permission is mandatory for trusted publishing
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install build
- name: Build package
run: python -m build
- name: Publish package to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
with:
password: ${{ secrets.PYPI_API_TOKEN }}
skip-existing: true
-215
View File
@@ -1,215 +0,0 @@
# Byte-compiled / optimized / DLL files
__pycache__/
*.py[cod]
*$py.class
# C extensions
*.so
# Distribution / packaging
.Python
build/
develop-eggs/
dist/
downloads/
eggs/
.eggs/
lib/
lib64/
parts/
sdist/
var/
wheels/
share/python-wheels/
*.egg-info/
.installed.cfg
*.egg
MANIFEST
# PyInstaller
# Usually these files are written by a python script from a template
# before PyInstaller builds the exe, so as to inject date/other infos into it.
*.manifest
*.spec
# Installer logs
pip-log.txt
pip-delete-this-directory.txt
# Unit test / coverage reports
htmlcov/
.tox/
.nox/
.coverage
.coverage.*
.cache
nosetests.xml
coverage.xml
*.cover
*.py,cover
.hypothesis/
.pytest_cache/
cover/
# Translations
*.mo
*.pot
# Django stuff:
*.log
local_settings.py
db.sqlite3
db.sqlite3-journal
# Flask stuff:
instance/
.webassets-cache
# Scrapy stuff:
.scrapy
# Sphinx documentation
docs/_build/
# PyBuilder
.pybuilder/
target/
# Jupyter Notebook
.ipynb_checkpoints
# IPython
profile_default/
ipython_config.py
# pyenv
# For a library or package, you might want to ignore these files since the code is
# intended to run in multiple environments; otherwise, check them in:
# .python-version
# pipenv
# According to pypa/pipenv#598, it is recommended to include Pipfile.lock in version control.
# However, in case of collaboration, if having platform-specific dependencies or dependencies
# having no cross-platform support, pipenv may install dependencies that don't work, or not
# install all needed dependencies.
#Pipfile.lock
# poetry
# Similar to Pipfile.lock, it is generally recommended to include poetry.lock in version control.
# This is especially recommended for binary packages to ensure reproducibility, and is more
# commonly ignored for libraries.
# https://python-poetry.org/docs/basic-usage/#commit-your-poetrylock-file-to-version-control
#poetry.lock
# pdm
# Similar to Pipfile.lock, it is generally recommended to include pdm.lock in version control.
#pdm.lock
# pdm stores project-wide configurations in .pdm.toml, but it is recommended to not include it
# in version control.
# https://pdm.fming.dev/#use-with-ide
.pdm.toml
# PEP 582; used by e.g. github.com/David-OConnor/pyflow and github.com/pdm-project/pdm
__pypackages__/
# Celery stuff
celerybeat-schedule
celerybeat.pid
# SageMath parsed files
*.sage.py
# Environments
.env
.venv
env/
venv/
ENV/
env.bak/
venv.bak/
# Spyder project settings
.spyderproject
.spyproject
# Rope project settings
.ropeproject
# mkdocs documentation
/site
# mypy
.mypy_cache/
.dmypy.json
dmypy.json
# Pyre type checker
.pyre/
# pytype static type analyzer
.pytype/
# Cython debug symbols
cython_debug/
# PyCharm
# JetBrains specific template is maintained in a separate JetBrains.gitignore that can
# be found at https://github.com/github/gitignore/blob/main/Global/JetBrains.gitignore
# and can be added to the global gitignore or merged into this file. For a more nuclear
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/
.DS_Store
hello.py
*.html
fast-mcp-docs.md
# Debug and test files
debug_*
test_*
CLAUDE.md
# ASGI/Deployment files
ssl/
*.pem
*.key
*.crt
# Docker volumes
redis-data/
# Production logs
logs/*.log.*
# Remove these lines - we need deployment files in git:
# Dockerfile - NEEDED for SaaS deployment
# fly.toml - NEEDED for Fly.io deployment
# .github/workflows/fly-deploy.yml - NEEDED for GitHub Actions
GEMINI.md
fly.toml
scripts/deploy-flyio.sh
docs/DEPLOYMENT_FLYIO.md
setup_jwt_template.py
mcp_server_main.py.backup
mcp_overhead_content.json
ANTHROPIC_TEST_README.md
extract_mcp_overhead.py
mcp_overhead_content.txt
mcp_overhead_summary.txt
run_http_server.py
run_local_test.py
# MCP overhead analysis files
mcp_overhead_*.json
mcp_overhead_*.txt
mcp_test_results_*.json
mcp_quick_test_*.json
# General text files (temporary notes, etc)
*.txt
analyze_playwright_mcp.py
measure_mcp_directly.py
playwright_mcp_overhead.json
simple_test.py
analyze_anayasa_html.py
Binary file not shown.

Before

Width:  |  Height:  |  Size: 73 KiB

-29
View File
@@ -1,29 +0,0 @@
# -------- BASE IMAGE (includes Chromium & deps) ----------------------------
FROM mcr.microsoft.com/playwright/python:v1.53.0-noble
# -------- Runtime setup ----------------------------------------------------
WORKDIR /app
# Copy dependency manifests first for layer-cache
COPY pyproject.toml poetry.lock* requirements*.txt* ./
# Fast, deterministic install with `uv`
RUN pip install --no-cache-dir uv && \
uv pip install --system --no-cache-dir .[asgi,saas]
# Copy application source
COPY . .
# -------- Environment ------------------------------------------------------
ENV PYTHONUNBUFFERED=1
ENV ENABLE_AUTH=true
ENV PORT=8000
# -------- Health check -----------------------------------------------------
HEALTHCHECK --interval=30s --timeout=10s --start-period=10s --retries=3 \
CMD python -c "import httpx, os, sys; r=httpx.get(f'http://localhost:{os.getenv(\"PORT\",\"8000\")}/health'); sys.exit(0 if r.status_code==200 else 1)"
EXPOSE 8000
# -------- Entrypoint -------------------------------------------------------
CMD ["uvicorn", "asgi_app:app", "--host", "0.0.0.0", "--port", "8000", "--proxy-headers"]
-21
View File
@@ -1,21 +0,0 @@
MIT License
Copyright (c) 2025 saidsurucu
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
-1
View File
@@ -1 +0,0 @@
web: uvicorn asgi_app:app --host 0.0.0.0 --port $PORT
-272
View File
@@ -1,272 +0,0 @@
# Yargı MCP: Türk Hukuk Kaynakları için MCP Sunucusu
[![Star History Chart](https://api.star-history.com/svg?repos=saidsurucu/yargi-mcp&type=Date)](https://www.star-history.com/#saidsurucu/yargi-mcp&Date)
Bu proje, çeşitli Türk hukuk kaynaklarına (Yargıtay, Danıştay, Emsal Kararlar, Uyuşmazlık Mahkemesi, Anayasa Mahkemesi - Norm Denetimi ile Bireysel Başvuru Kararları, Kamu İhale Kurulu Kararları, Rekabet Kurumu Kararları, Sayıştay Kararları, KVKK Kararları ve BDDK Kararları) erişimi kolaylaştıran bir [FastMCP](https://gofastmcp.com/) sunucusu oluşturur. Bu sayede, bu kaynaklardan veri arama ve belge getirme işlemleri, Model Context Protocol (MCP) destekleyen LLM (Büyük Dil Modeli) uygulamaları (örneğin Claude Desktop veya [5ire](https://5ire.app)) ve diğer istemciler tarafından araç (tool) olarak kullanılabilir hale gelir.
![örnek](./ornek.png)
🎯 **Temel Özellikler**
🚀 **YÜKSEK PERFORMANS OPTİMİZASYONU:** Bu MCP sunucusu **%61.8 token azaltma** ile optimize edilmiştir (8,692 token tasarrufu). Claude AI ile daha hızlı yanıt süreleri ve daha verimli etkileşim sağlar.
* Çeşitli Türk hukuk veritabanlarına programatik erişim için standart bir MCP arayüzü.
* **Kapsamlı Mahkeme Daire/Kurul Filtreleme:** 79 farklı daire/kurul filtreleme seçeneği
* **Dual/Triple API Desteği:** Her mahkeme için birden fazla API kaynağı ile maksimum kapsama
* **Kapsamlı Tarih Filtreleme:** Tüm Bedesten API araçlarında ISO 8601 formatında tarih aralığı filtreleme
* **Kesin Cümle Arama:** Tüm Bedesten API araçlarında çift tırnak ile tam cümle arama desteği
* Aşağıdaki kurumların kararlarını arama ve getirme yeteneği:
* **Yargıtay:** Detaylı kriterlerle karar arama ve karar metinlerini Markdown formatında getirme. **Dual API** (Ana + Bedesten) + **52 Daire/Kurul Filtreleme** + **Tarih & Kesin Cümle Arama** (Hukuk/Ceza Daireleri, Genel Kurullar)
* **Danıştay:** Anahtar kelime bazlı ve detaylı kriterlerle karar arama; karar metinlerini Markdown formatında getirme. **Triple API** (Keyword + Detailed + Bedesten) + **27 Daire/Kurul Filtreleme** + **Tarih & Kesin Cümle Arama** (İdari Daireler, Vergi/İdare Kurulları, Askeri Yüksek İdare Mahkemesi)
* **Yerel Hukuk Mahkemeleri:** Bedesten API ile yerel hukuk mahkemesi kararlarına erişim + **Tarih & Kesin Cümle Arama**
* **İstinaf Hukuk Mahkemeleri:** Bedesten API ile istinaf mahkemesi kararlarına erişim + **Tarih & Kesin Cümle Arama**
* **Kanun Yararına Bozma (KYB):** Bedesten API ile olağanüstü kanun yoluna erişim + **Tarih & Kesin Cümle Arama**
* **Emsal (UYAP):** Detaylı kriterlerle emsal karar arama ve karar metinlerini Markdown formatında getirme.
* **Uyuşmazlık Mahkemesi:** Form tabanlı kriterlerle karar arama ve karar metinlerini (URL ile erişilen) Markdown formatında getirme.
* **Anayasa Mahkemesi (Norm Denetimi):** Kapsamlı kriterlerle norm denetimi kararlarını arama; uzun karar metinlerini (5.000 karakterlik) sayfalanmış Markdown formatında getirme.
* **Anayasa Mahkemesi (Bireysel Başvuru):** Kapsamlı kriterlerle bireysel başvuru "Karar Arama Raporu" oluşturma ve listedeki kararların metinlerini (5.000 karakterlik) sayfalanmış Markdown formatında getirme.
* **KİK (Kamu İhale Kurulu):** Çeşitli kriterlerle Kurul kararlarını arama; uzun karar metinlerini (varsayılan 5.000 karakterlik) sayfalanmış Markdown formatında getirme.
* **Rekabet Kurumu:** Çeşitli kriterlerle Kurul kararlarını arama; karar metinlerini Markdown formatında getirme.
* **Sayıştay:** 3 karar türü ile kapsamlı denetim kararlarına erişim + **8 Daire Filtreleme** + **Tarih Aralığı & İçerik Arama** (Genel Kurul yorumlayıcı kararları, Temyiz Kurulu itiraz kararları, Daire ilk derece denetim kararları)
* **KVKK (Kişisel Verilerin Korunması Kurulu):** Brave Search API ile veri koruma kararlarını arama; uzun karar metinlerini (5.000 karakterlik) sayfalanmış Markdown formatında getirme + **Türkçe Arama** + **Site Hedeflemeli Arama** (kvkk.gov.tr kararları)
* **BDDK (Bankacılık Düzenleme ve Denetleme Kurumu):** Bankacılık düzenleme kararlarını arama; karar metinlerini Markdown formatında getirme + **Optimized Search** + **"Karar Sayısı" Targeting** + **Spesifik URL Filtreleme** (bddk.org.tr/Mevzuat/DokumanGetir)
* Karar metinlerinin daha kolay işlenebilmesi için Markdown formatına çevrilmesi.
* Claude Desktop uygulaması ile `fastmcp install` komutu kullanılarak kolay entegrasyon.
* Yargı MCP artık [5ire](https://5ire.app) gibi Claude Desktop haricindeki MCP istemcilerini de destekliyor!
---
<details>
<summary>🚀 <strong>Claude Haricindeki Modellerle Kullanmak İçin Çok Kolay Kurulum (Örnek: 5ire için)</strong></summary>
Bu bölüm, Yargı MCP aracını 5ire gibi Claude Desktop dışındaki MCP istemcileriyle kullanmak isteyenler içindir.
* **Python Kurulumu:** Sisteminizde Python 3.11 veya üzeri kurulu olmalıdır. Kurulum sırasında "**Add Python to PATH**" (Python'ı PATH'e ekle) seçeneğini işaretlemeyi unutmayın. [Buradan](https://www.python.org/downloads/) indirebilirsiniz.
* **Git Kurulumu (Windows):** Bilgisayarınıza [git](https://git-scm.com/downloads/win) yazılımını indirip kurun. "Git for Windows/x64 Setup" seçeneğini indirmelisiniz.
* **`uv` Kurulumu:**
* **Windows Kullanıcıları (PowerShell):** Bir CMD ekranı açın ve bu kodu çalıştırın: `powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"`
* **Mac/Linux Kullanıcıları (Terminal):** Bir Terminal ekranı açın ve bu kodu çalıştırın: `curl -LsSf https://astral.sh/uv/install.sh | sh`
* **Microsoft Visual C++ Redistributable (Windows):** Bazı Python paketlerinin doğru çalışması için gereklidir. [Buradan](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170) indirip kurun.
* İşletim sisteminize uygun [5ire](https://5ire.app) MCP istemcisini indirip kurun.
* 5ire'ı açın. **Workspace -> Providers** menüsünden kullanmak istediğiniz LLM servisinin API anahtarını girin.
* **Tools** menüsüne girin. **+Local** veya **New** yazan butona basın.
* **Tool Key:** `yargimcp`
* **Name:** `Yargı MCP`
* **Command:**
```
uvx yargi-mcp
```
* **Save** butonuna basarak kaydedin.
![5ire ayarları](./5ire-settings.png)
* Şimdi **Tools** altında **Yargı MCP**'yi görüyor olmalısınız. Üstüne geldiğinizde sağda çıkan butona tıklayıp etkinleştirin (yeşil ışık yanmalı).
* Artık Yargı MCP ile konuşabilirsiniz.
</details>
---
<details>
<summary>⚙️ <strong>Claude Desktop Manuel Kurulumu</strong></summary>
1. **Ön Gereksinimler:** Python, `uv`, (Windows için) Microsoft Visual C++ Redistributable'ın sisteminizde kurulu olduğundan emin olun. Detaylı bilgi için yukarıdaki "5ire için Kurulum" bölümündeki ilgili adımlara bakabilirsiniz.
2. Claude Desktop **Settings -> Developer -> Edit Config**.
3. Açılan `claude_desktop_config.json` dosyasına `mcpServers` altına ekleyin:
```json
{
"mcpServers": {
// ... (varsa diğer sunucularınız) ...
"Yargı MCP": {
"command": "uvx",
"args": [
"yargi-mcp"
]
}
}
}
```
4. Claude Desktop'ı kapatıp yeniden başlatın.
</details>
---
<details>
<summary>🌟 <strong>Gemini CLI ile Kullanım</strong></summary>
Yargı MCP'yi Gemini CLI ile kullanmak için:
1. **Ön Gereksinimler:** Python, `uv`, (Windows için) Microsoft Visual C++ Redistributable'ın sisteminizde kurulu olduğundan emin olun. Detaylı bilgi için yukarıdaki "5ire için Kurulum" bölümündeki ilgili adımlara bakabilirsiniz.
2. **Gemini CLI ayarlarını yapılandırın:**
Gemini CLI'ın ayar dosyasını düzenleyin:
- **macOS/Linux:** `~/.gemini/settings.json`
- **Windows:** `%USERPROFILE%\.gemini\settings.json`
Aşağıdaki `mcpServers` bloğunu ekleyin:
```json
{
"theme": "Default",
"selectedAuthType": "###",
"mcpServers": {
"yargi_mcp": {
"command": "uvx",
"args": [
"yargi-mcp"
]
}
}
}
```
**Yapılandırma açıklamaları:**
- `"yargi_mcp"`: Sunucunuz için yerel bir isim
- `"command"`: `uvx` komutu (uv'nin paket çalıştırma aracı)
- `"args"`: GitHub'dan doğrudan Yargı MCP'yi çalıştırmak için gerekli argümanlar
3. **Kullanım:**
- Gemini CLI'ı başlatın
- Yargı MCP araçları otomatik olarak kullanılabilir olacaktır
- Örnek komutlar:
- "Yargıtay'ın mülkiyet hakkı ile ilgili son kararlarını ara"
- "Danıştay'ın imar planı iptaline ilişkin kararlarını bul"
- "Anayasa Mahkemesi'nin ifade özgürlüğü kararlarını getir"
</details>
<details>
<summary>🛠️ <strong>Kullanılabilir Araçlar (MCP Tools)</strong></summary>
Bu FastMCP sunucusu **19 optimize edilmiş MCP aracı** sunar (token verimliliği için optimize edilmiş):
### **Yargıtay Araçları (Birleşik Bedesten API - Token Optimized)**
*Not: Yargıtay araçları token verimliliği için birleşik Bedesten API'ye entegre edilmiştir*
### **Danıştay Araçları (Birleşik Bedesten API - Token Optimized)**
*Not: Danıştay araçları token verimliliği için birleşik Bedesten API'ye entegre edilmiştir*
### **Birleşik Bedesten API Araçları (5 Mahkeme) - 🚀 TOKEN OPTİMİZE**
1. `search_bedesten_unified(phrase, court_types, birimAdi, kararTarihiStart, kararTarihiEnd, ...)`: **5 mahkeme türünü** birleşik arama (Yargıtay, Danıştay, Yerel Hukuk, İstinaf Hukuk, KYB) + **79 daire filtreleme** + **Tarih & Kesin Cümle Arama**
2. `get_bedesten_document_markdown(documentId: str)`: Bedesten API'den herhangi bir belgeyi Markdown formatında getirir (HTML/PDF → Markdown)
### **Emsal Karar Araçları (UYAP)**
3. `search_emsal_detailed_decisions(keyword, ...)`: Emsal (UYAP) kararlarını detaylı kriterlerle arar.
4. `get_emsal_document_markdown(id: str)`: Belirli bir Emsal kararının metnini Markdown formatında getirir.
### **Uyuşmazlık Mahkemesi Araçları**
5. `search_uyusmazlik_decisions(icerik, ...)`: Uyuşmazlık Mahkemesi kararlarını çeşitli form kriterleriyle arar.
6. `get_uyusmazlik_document_markdown_from_url(document_url)`: Bir Uyuşmazlık kararını tam URL'sinden alıp Markdown formatında getirir.
### **Anayasa Mahkemesi Araçları (Birleşik API) - 🚀 TOKEN OPTİMİZE**
7. `search_anayasa_unified(decision_type, keywords_all, ...)`: AYM kararlarını birleşik arama (Norm Denetimi + Bireysel Başvuru) - **4 araç → 2 araç optimizasyonu**
8. `get_anayasa_document_unified(document_url, page_number)`: AYM kararlarını birleşik belge getirme - **sayfalanmış Markdown** içeriği
### **KİK (Kamu İhale Kurulu) Araçları**
9. `search_kik_decisions(karar_tipi, ...)`: KİK (Kamu İhale Kurulu) kararlarını arar.
10. `get_kik_document_markdown(karar_id, page_number)`: Belirli bir KİK kararını, Base64 ile encode edilmiş `karar_id`'sini kullanarak alır ve **sayfalanmış Markdown** içeriğini getirir.
### **Rekabet Kurumu Araçları**
    * `search_rekabet_kurumu_decisions(KararTuru: Literal[...], ...) -> RekabetSearchResult`: Rekabet Kurumu kararlarını arar. `KararTuru` için kullanıcı dostu isimler kullanılır (örn: "Birleşme ve Devralma").
    * `get_rekabet_kurumu_document(karar_id: str, page_number: Optional[int] = 1) -> RekabetDocument`: Belirli bir Rekabet Kurumu kararını `karar_id` ile alır. Kararın PDF formatındaki orijinalinden istenen sayfayı ayıklar ve Markdown formatında döndürür.
---
* **Sayıştay Araçları (3 Karar Türü + 8 Daire Filtreleme):**
* `search_sayistay_genel_kurul(karar_no, karar_tarih_baslangic, karar_tamami, ...)`: Sayıştay Genel Kurul (yorumlayıcı) kararlarını arar. **Tarih aralığı** (2006-2024) + **İçerik arama** (400 karakter)
* `search_sayistay_temyiz_kurulu(ilam_dairesi, kamu_idaresi_turu, temyiz_karar, ...)`: Temyiz Kurulu (itiraz) kararlarını arar. **8 Daire filtreleme** + **Kurum türü** + **Konu sınıflandırması**
* `search_sayistay_daire(yargilama_dairesi, web_karar_metni, hesap_yili, ...)`: Daire (ilk derece denetim) kararlarını arar. **8 Daire filtreleme** + **Hesap yılı** + **İçerik arama**
* `get_sayistay_genel_kurul_document_markdown(decision_id: str)`: Genel Kurul kararının tam metnini Markdown formatında getirir
* `get_sayistay_temyiz_kurulu_document_markdown(decision_id: str)`: Temyiz Kurulu kararının tam metnini Markdown formatında getirir
* `get_sayistay_daire_document_markdown(decision_id: str)`: Daire kararının tam metnini Markdown formatında getirir
* **KVKK Araçları (Brave Search API + Türkçe Arama):**
* `search_kvkk_decisions(keywords, page, pageSize, ...)`: KVKK (Kişisel Verilerin Korunması Kurulu) kararlarını Brave Search API ile arar. **Türkçe arama** + **Site hedeflemeli** (`site:kvkk.gov.tr "karar özeti"`) + **Sayfalama desteği**
* `get_kvkk_document_markdown(decision_url: str, page_number: Optional[int] = 1)`: KVKK kararının tam metnini **sayfalanmış Markdown** formatında getirir (5.000 karakterlik sayfa)
### BDDK Araçları
* `search_bddk_decisions(keywords, page)`: BDDK (Bankacılık Düzenleme ve Denetleme Kurumu) kararlarını arar. **"Karar Sayısı" targeting** + **Spesifik URL filtreleme** (`bddk.org.tr/Mevzuat/DokumanGetir`) + **Optimized search**
* `get_bddk_document_markdown(document_id: str, page_number: Optional[int] = 1)`: BDDK kararının tam metnini **sayfalanmış Markdown** formatında getirir (5.000 karakterlik sayfa)
</details>
---
<details>
<summary>📊 <strong>Kapsamlı İstatistikler & Optimizasyon Başarıları</strong></summary>
🚀 **TOKEN OPTİMİZASYON BAŞARISI:**
- **%61.8 Token Azaltma:** 14,061 → 5,369 tokens (8,692 token tasarrufu)
- **Hedef Aşım:** 10,000 token hedefini 4,631 token aştık
- **Daha Hızlı Yanıt:** Claude AI ile optimize edilmiş etkileşim
- **Korunan İşlevsellik:** %100 özellik desteği devam ediyor
**GENEL İSTATİSTİKLER:**
- **Toplam Mahkeme/Kurum:** 13 farklı hukuki kurum (KVKK dahil)
- **Toplam MCP Tool:** 19 optimize edilmiş arama ve belge getirme aracı
- **Daire/Kurul Filtreleme:** 87 farklı seçenek (52 Yargıtay + 27 Danıştay + 8 Sayıştay)
- **Tarih Filtreleme:** Birleşik Bedesten API aracında ISO 8601 formatında tam tarih aralığı desteği
- **Kesin Cümle Arama:** Birleşik Bedesten API aracında çift tırnak ile tam cümle arama (`"\"mülkiyet kararı\""` formatı)
- **Birleşik API:** 10 ayrı Bedesten aracı → 2 birleşik araç (search_bedesten_unified + get_bedesten_document_markdown)
- **API Kaynağı:** Dual/Triple API desteği ile maksimum kapsama
- **Tam Türk Adalet Sistemi:** Yerel mahkemelerden en yüksek mahkemelere kadar
**🏛️ Desteklenen Mahkeme Hiyerarşisi:**
```
Yerel Mahkemeler → İstinaf → Yargıtay/Danıştay → Anayasa Mahkemesi
↓ ↓ ↓ ↓
Bedesten API Bedesten API Dual/Triple API Norm+Bireysel API
+ Tarih + Kesin + Tarih + Kesin + Daire + Tarih + Gelişmiş
Cümle Arama Cümle Arama + Kesin Cümle Arama
```
**⚖️ Kapsamlı Filtreleme Özellikleri:**
- **Daire Filtreleme:** 79 seçenek (52 Yargıtay + 27 Danıştay)
- **Yargıtay:** 52 seçenek (1-23 Hukuk, 1-23 Ceza, Genel Kurullar, Başkanlar Kurulu)
- **Danıştay:** 27 seçenek (1-17 Daireler, İdare/Vergi Kurulları, Askeri Mahkemeler)
- **Tarih Filtreleme:** 5 Bedesten API aracında ISO 8601 formatı (YYYY-MM-DDTHH:MM:SS.000Z)
- Tek tarih, tarih aralığı, tek taraflı filtreleme desteği
- Yargıtay, Danıştay, Yerel Hukuk, İstinaf Hukuk, KYB kararları
- **Kesin Cümle Arama:** 5 Bedesten API aracında çift tırnak formatı
- Normal arama: `"mülkiyet kararı"` (kelimeler ayrı ayrı)
- Kesin arama: `"\"mülkiyet kararı\""` (tam cümle olarak)
- Daha kesin sonuçlar için hukuki terimler ve kavramlar
**🔧 OPTİMİZASYON DETAYLARI:**
- **Anayasa Mahkemesi:** 4 araç → 2 birleşik araç (search_anayasa_unified + get_anayasa_document_unified)
- **Yargıtay & Danıştay:** Ana API araçları birleşik Bedesten API'ye entegre edildi
- **Sayıştay:** 6 araç → 2 birleşik araç (search_sayistay_unified + get_sayistay_document_unified)
- **Parameter Optimizasyonu:** pageSize parametreleri optimize edildi
- **Açıklama Optimizasyonu:** Uzun açıklamalar kısaltıldı (örn: KIK karar_metni)
</details>
---
<details>
<summary>🌐 <strong>Web Service / ASGI Deployment</strong></summary>
Yargı MCP artık web servisi olarak da çalıştırılabilir! ASGI desteği sayesinde:
- **Web API olarak erişim**: HTTP endpoint'leri üzerinden MCP araçlarına erişim
- **Cloud deployment**: Heroku, Railway, Google Cloud Run, AWS Lambda desteği
- **Docker desteği**: Production-ready Docker container
- **FastAPI entegrasyonu**: REST API ve interaktif dokümantasyon
**Hızlı başlangıç:**
```bash
# ASGI dependencies yükle
pip install yargi-mcp[asgi]
# Web servisi olarak başlat
python run_asgi.py
# veya
uvicorn asgi_app:app --host 0.0.0.0 --port 8000
```
Detaylı deployment rehberi için: [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md)
</details>
---
📜 **Lisans**
Bu proje MIT Lisansı altında lisanslanmıştır. Detaylar için `LICENSE` dosyasına bakınız.
-7
View File
@@ -1,7 +0,0 @@
#!/usr/bin/env python3
"""Entry point for yargi-mcp package."""
from mcp_server_main import main
if __name__ == "__main__":
main()
@@ -1,355 +0,0 @@
# anayasa_mcp_module/bireysel_client.py
# This client is for Bireysel Başvuru: https://kararlarbilgibankasi.anayasa.gov.tr
import httpx
from bs4 import BeautifulSoup, Tag
from typing import Dict, Any, List, Optional, Tuple
import logging
import html
import re
import io
from urllib.parse import urlencode, urljoin, quote
from markitdown import MarkItDown
import math # For math.ceil for pagination
from .models import (
AnayasaBireyselReportSearchRequest,
AnayasaBireyselReportDecisionDetail,
AnayasaBireyselReportDecisionSummary,
AnayasaBireyselReportSearchResult,
AnayasaBireyselBasvuruDocumentMarkdown, # Model for Bireysel Başvuru document
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
class AnayasaBireyselBasvuruApiClient:
BASE_URL = "https://kararlarbilgibankasi.anayasa.gov.tr"
SEARCH_PATH = "/Ara"
DOCUMENT_MARKDOWN_CHUNK_SIZE = 5000 # Character limit per page
def __init__(self, request_timeout: float = 60.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
},
timeout=request_timeout,
verify=True,
follow_redirects=True
)
def _build_query_params_for_bireysel_report(self, params: AnayasaBireyselReportSearchRequest) -> List[Tuple[str, str]]:
query_params: List[Tuple[str, str]] = []
query_params.append(("KararBulteni", "1")) # Specific to this report type
if params.keywords:
for kw in params.keywords:
query_params.append(("KelimeAra[]", kw))
if params.page_to_fetch and params.page_to_fetch > 1:
query_params.append(("page", str(params.page_to_fetch)))
return query_params
async def search_bireysel_basvuru_report(
self,
params: AnayasaBireyselReportSearchRequest
) -> AnayasaBireyselReportSearchResult:
final_query_params = self._build_query_params_for_bireysel_report(params)
request_url = self.SEARCH_PATH
logger.info(f"AnayasaBireyselBasvuruApiClient: Performing Bireysel Başvuru Report search. Path: {request_url}, Params: {final_query_params}")
try:
response = await self.http_client.get(request_url, params=final_query_params)
response.raise_for_status()
html_content = response.text
except httpx.RequestError as e:
logger.error(f"AnayasaBireyselBasvuruApiClient: HTTP request error during Bireysel Başvuru Report search: {e}")
raise
except Exception as e:
logger.error(f"AnayasaBireyselBasvuruApiClient: Error processing Bireysel Başvuru Report search request: {e}")
raise
soup = BeautifulSoup(html_content, 'html.parser')
total_records = None
bulunan_karar_div = soup.find("div", class_="bulunankararsayisi")
if bulunan_karar_div:
match_records = re.search(r'(\d+)\s*Karar Bulundu', bulunan_karar_div.get_text(strip=True))
if match_records:
total_records = int(match_records.group(1))
processed_decisions: List[AnayasaBireyselReportDecisionSummary] = []
report_content_area = soup.find("div", class_="HaberBulteni")
if not report_content_area:
logger.warning("HaberBulteni div not found, attempting to parse decision divs from the whole page.")
report_content_area = soup
decision_divs = report_content_area.find_all("div", class_="KararBulteniBirKarar")
if not decision_divs:
logger.warning("No KararBulteniBirKarar divs found.")
for decision_div in decision_divs:
title_tag = decision_div.find("h4")
title_text = title_tag.get_text(strip=True) if title_tag and title_tag.strong else (title_tag.get_text(strip=True) if title_tag else "")
alti_cizili_div = decision_div.find("div", class_="AltiCizili")
ref_no, dec_type, body, app_date, dec_date, url_path = "", "", "", "", "", ""
if alti_cizili_div:
link_tag = alti_cizili_div.find("a", href=True)
if link_tag:
ref_no = link_tag.get_text(strip=True)
url_path = link_tag['href']
parts_text = alti_cizili_div.get_text(separator="|", strip=True)
parts = [part.strip() for part in parts_text.split("|")]
# Clean ref_no from the first part if it was extracted from link
if ref_no and parts and parts[0].strip().startswith(ref_no):
parts[0] = parts[0].replace(ref_no, "").strip()
if not parts[0]: parts.pop(0) # Remove empty string if ref_no was the only content
# Assign parts based on typical order, adjusting for missing ref_no at start
current_idx = 0
if not ref_no and len(parts) > current_idx and re.match(r"\d+/\d+", parts[current_idx]): # Check if first part is ref_no
ref_no = parts[current_idx]
current_idx += 1
dec_type = parts[current_idx] if len(parts) > current_idx else ""
current_idx += 1
body = parts[current_idx] if len(parts) > current_idx else ""
current_idx += 1
app_date_raw = parts[current_idx] if len(parts) > current_idx else ""
current_idx += 1
dec_date_raw = parts[current_idx] if len(parts) > current_idx else ""
if app_date_raw and "Başvuru Tarihi :" in app_date_raw:
app_date = app_date_raw.replace("Başvuru Tarihi :", "").strip()
elif app_date_raw: # If label is missing but format matches
app_date_match = re.search(r'(\d{1,2}/\d{1,2}/\d{4})', app_date_raw)
if app_date_match: app_date = app_date_match.group(1)
if dec_date_raw and "Karar Tarihi :" in dec_date_raw:
dec_date = dec_date_raw.replace("Karar Tarihi :", "").strip()
elif dec_date_raw: # If label is missing but format matches
dec_date_match = re.search(r'(\d{1,2}/\d{1,2}/\d{4})', dec_date_raw)
if dec_date_match: dec_date = dec_date_match.group(1)
subject_div = decision_div.find(lambda tag: tag.name == 'div' and not tag.has_attr('class') and tag.get_text(strip=True).startswith("BAŞVURU KONUSU :"))
subject_text = subject_div.get_text(strip=True).replace("BAŞVURU KONUSU :", "").strip() if subject_div else ""
details_list: List[AnayasaBireyselReportDecisionDetail] = []
karar_detaylari_div = decision_div.find_next_sibling("div", id="KararDetaylari") # Corrected: was KararDetaylari
if karar_detaylari_div:
table = karar_detaylari_div.find("table", class_="table")
if table and table.find("tbody"):
for row in table.find("tbody").find_all("tr"):
cells = row.find_all("td")
if len(cells) == 4: # Hak, Müdahale İddiası, Sonuç, Giderim
details_list.append(AnayasaBireyselReportDecisionDetail(
hak=cells[0].get_text(strip=True) or "",
mudahale_iddiasi=cells[1].get_text(strip=True) or "",
sonuc=cells[2].get_text(strip=True) or "",
giderim=cells[3].get_text(strip=True) or "",
))
full_decision_page_url = urljoin(self.BASE_URL, url_path) if url_path else ""
processed_decisions.append(AnayasaBireyselReportDecisionSummary(
title=title_text,
decision_reference_no=ref_no,
decision_page_url=full_decision_page_url,
decision_type_summary=dec_type,
decision_making_body=body,
application_date_summary=app_date,
decision_date_summary=dec_date,
application_subject_summary=subject_text,
details=details_list
))
return AnayasaBireyselReportSearchResult(
decisions=processed_decisions,
total_records_found=total_records,
retrieved_page_number=params.page_to_fetch
)
def _convert_html_to_markdown_bireysel(self, full_decision_html_content: str) -> Optional[str]:
if not full_decision_html_content:
return None
processed_html = html.unescape(full_decision_html_content)
soup = BeautifulSoup(processed_html, "html.parser")
html_input_for_markdown = ""
karar_tab_content = soup.find("div", id="Karar")
if karar_tab_content:
karar_html_span = karar_tab_content.find("span", class_="kararHtml")
if karar_html_span:
word_section = karar_html_span.find("div", class_="WordSection1")
if word_section:
for s in word_section.select('script, style, .item.col-xs-12.col-sm-12, center:has(b)'):
s.decompose()
html_input_for_markdown = str(word_section)
else:
logger.warning("AnayasaBireyselBasvuruApiClient: WordSection1 not found in span.kararHtml. Using span.kararHtml content.")
for s in karar_html_span.select('script, style, .item.col-xs-12.col-sm-12, center:has(b)'):
s.decompose()
html_input_for_markdown = str(karar_html_span)
else:
logger.warning("AnayasaBireyselBasvuruApiClient: span.kararHtml not found in div#Karar. Using div#Karar content.")
for s in karar_tab_content.select('script, style, .item.col-xs-12.col-sm-12, center:has(b)'):
s.decompose()
html_input_for_markdown = str(karar_tab_content)
else:
logger.warning("AnayasaBireyselBasvuruApiClient: div#Karar (KARAR tab) not found. Trying WordSection1 fallback.")
word_section_fallback = soup.find("div", class_="WordSection1")
if word_section_fallback:
for s in word_section_fallback.select('script, style, .item.col-xs-12.col-sm-12, center:has(b)'):
s.decompose()
html_input_for_markdown = str(word_section_fallback)
else:
body_tag = soup.find("body")
if body_tag:
for s in body_tag.select('script, style, .item.col-xs-12.col-sm-12, center:has(b), .banner, .footer, .yazdirmaalani, .filtreler, .menu, .altmenu, .geri, .arabuton, .temizlebutonu, form#KararGetir, .TabBaslik, #KararDetaylari, .share-button-container'):
s.decompose()
html_input_for_markdown = str(body_tag)
else:
html_input_for_markdown = processed_html
markdown_text = None
try:
# Ensure the content is wrapped in basic HTML structure if it's not already
if not html_input_for_markdown.strip().lower().startswith(("<html", "<!doctype")):
html_content = f"<html><head><meta charset=\"UTF-8\"></head><body>{html_input_for_markdown}</body></html>"
else:
html_content = html_input_for_markdown
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_content.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_text = conversion_result.text_content
except Exception as e:
logger.error(f"AnayasaBireyselBasvuruApiClient: MarkItDown conversion error: {e}")
return markdown_text
async def get_decision_document_as_markdown(
self,
document_url_path: str, # e.g. /BB/2021/20295
page_number: int = 1
) -> AnayasaBireyselBasvuruDocumentMarkdown:
full_url = urljoin(self.BASE_URL, document_url_path)
logger.info(f"AnayasaBireyselBasvuruApiClient: Fetching Bireysel Başvuru document for Markdown (page {page_number}) from URL: {full_url}")
basvuru_no_from_page = None
karar_tarihi_from_page = None
basvuru_tarihi_from_page = None
karari_veren_birim_from_page = None
karar_turu_from_page = None
resmi_gazete_info_from_page = None
try:
response = await self.http_client.get(full_url)
response.raise_for_status()
html_content_from_api = response.text
if not isinstance(html_content_from_api, str) or not html_content_from_api.strip():
logger.warning(f"AnayasaBireyselBasvuruApiClient: Received empty HTML from {full_url}.")
return AnayasaBireyselBasvuruDocumentMarkdown(
source_url=full_url, markdown_chunk=None, current_page=page_number, total_pages=0, is_paginated=False
)
soup = BeautifulSoup(html_content_from_api, 'html.parser')
meta_desc_tag = soup.find("meta", attrs={"name": "description"})
if meta_desc_tag and meta_desc_tag.get("content"):
content = meta_desc_tag["content"]
bn_match = re.search(r"B\.\s*No:\s*([\d\/]+)", content)
if bn_match: basvuru_no_from_page = bn_match.group(1).strip()
date_match = re.search(r"(\d{1,2}\/\d{1,2}\/\d{4}),\s*§", content)
if date_match: karar_tarihi_from_page = date_match.group(1).strip()
karar_detaylari_tab = soup.find("div", id="KararDetaylari")
if karar_detaylari_tab:
table = karar_detaylari_tab.find("table", class_="table")
if table:
rows = table.find_all("tr")
for row in rows:
cells = row.find_all("td")
if len(cells) == 2:
key = cells[0].get_text(strip=True)
value = cells[1].get_text(strip=True)
if "Kararı Veren Birim" in key: karari_veren_birim_from_page = value
elif "Karar Türü (Başvuru Sonucu)" in key: karar_turu_from_page = value
elif "Başvuru No" in key and not basvuru_no_from_page: basvuru_no_from_page = value
elif "Başvuru Tarihi" in key: basvuru_tarihi_from_page = value
elif "Karar Tarihi" in key and not karar_tarihi_from_page: karar_tarihi_from_page = value
elif "Resmi Gazete Tarih / Sayı" in key: resmi_gazete_info_from_page = value
full_markdown_content = self._convert_html_to_markdown_bireysel(html_content_from_api)
if not full_markdown_content:
return AnayasaBireyselBasvuruDocumentMarkdown(
source_url=full_url,
basvuru_no_from_page=basvuru_no_from_page,
karar_tarihi_from_page=karar_tarihi_from_page,
basvuru_tarihi_from_page=basvuru_tarihi_from_page,
karari_veren_birim_from_page=karari_veren_birim_from_page,
karar_turu_from_page=karar_turu_from_page,
resmi_gazete_info_from_page=resmi_gazete_info_from_page,
markdown_chunk=None,
current_page=page_number,
total_pages=0,
is_paginated=False
)
content_length = len(full_markdown_content)
total_pages = math.ceil(content_length / self.DOCUMENT_MARKDOWN_CHUNK_SIZE)
if total_pages == 0: total_pages = 1
current_page_clamped = max(1, min(page_number, total_pages))
start_index = (current_page_clamped - 1) * self.DOCUMENT_MARKDOWN_CHUNK_SIZE
end_index = start_index + self.DOCUMENT_MARKDOWN_CHUNK_SIZE
markdown_chunk = full_markdown_content[start_index:end_index]
return AnayasaBireyselBasvuruDocumentMarkdown(
source_url=full_url,
basvuru_no_from_page=basvuru_no_from_page,
karar_tarihi_from_page=karar_tarihi_from_page,
basvuru_tarihi_from_page=basvuru_tarihi_from_page,
karari_veren_birim_from_page=karari_veren_birim_from_page,
karar_turu_from_page=karar_turu_from_page,
resmi_gazete_info_from_page=resmi_gazete_info_from_page,
markdown_chunk=markdown_chunk,
current_page=current_page_clamped,
total_pages=total_pages,
is_paginated=(total_pages > 1)
)
except httpx.RequestError as e:
logger.error(f"AnayasaBireyselBasvuruApiClient: HTTP error fetching Bireysel Başvuru document from {full_url}: {e}")
raise
except Exception as e:
logger.error(f"AnayasaBireyselBasvuruApiClient: General error processing Bireysel Başvuru document from {full_url}: {e}")
raise
async def close_client_session(self):
if hasattr(self, 'http_client') and self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("AnayasaBireyselBasvuruApiClient: HTTP client session closed.")
@@ -1,356 +0,0 @@
# anayasa_mcp_module/client.py
# This client is for Norm Denetimi: https://normkararlarbilgibankasi.anayasa.gov.tr
import httpx
from bs4 import BeautifulSoup
from typing import Dict, Any, List, Optional, Tuple
import logging
import html
import re
import io
from urllib.parse import urlencode, urljoin, quote
from markitdown import MarkItDown
import math # For math.ceil for pagination
from .models import (
AnayasaNormDenetimiSearchRequest,
AnayasaDecisionSummary,
AnayasaReviewedNormInfo,
AnayasaSearchResult,
AnayasaDocumentMarkdown, # Model for Norm Denetimi document
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
class AnayasaMahkemesiApiClient:
BASE_URL = "https://normkararlarbilgibankasi.anayasa.gov.tr"
SEARCH_PATH_SEGMENT = "Ara"
DOCUMENT_MARKDOWN_CHUNK_SIZE = 5000 # Character limit per page
def __init__(self, request_timeout: float = 60.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
},
timeout=request_timeout,
verify=True,
follow_redirects=True
)
def _build_search_query_params_for_aym(self, params: AnayasaNormDenetimiSearchRequest) -> List[Tuple[str, str]]:
query_params: List[Tuple[str, str]] = []
if params.keywords_all:
for kw in params.keywords_all: query_params.append(("KelimeAra[]", kw))
if params.keywords_any:
for kw in params.keywords_any: query_params.append(("HerhangiBirKelimeAra[]", kw))
if params.keywords_exclude:
for kw in params.keywords_exclude: query_params.append(("BulunmayanKelimeAra[]", kw))
if params.period and params.period and params.period != "ALL": query_params.append(("Donemler_id", params.period))
if params.case_number_esas: query_params.append(("EsasNo", params.case_number_esas))
if params.decision_number_karar: query_params.append(("KararNo", params.decision_number_karar))
if params.first_review_date_start: query_params.append(("IlkIncelemeTarihiIlk", params.first_review_date_start))
if params.first_review_date_end: query_params.append(("IlkIncelemeTarihiSon", params.first_review_date_end))
if params.decision_date_start: query_params.append(("KararTarihiIlk", params.decision_date_start))
if params.decision_date_end: query_params.append(("KararTarihiSon", params.decision_date_end))
if params.application_type and params.application_type and params.application_type != "ALL": query_params.append(("BasvuruTurler_id", params.application_type))
if params.applicant_general_name: query_params.append(("BasvuranGeneller_id", params.applicant_general_name))
if params.applicant_specific_name: query_params.append(("BasvuranOzeller_id", params.applicant_specific_name))
if params.attending_members_names:
for name in params.attending_members_names: query_params.append(("Uyeler_id[]", name))
if params.rapporteur_name: query_params.append(("Raportorler_id", params.rapporteur_name))
if params.norm_type and params.norm_type and params.norm_type != "ALL": query_params.append(("NormunTurler_id", params.norm_type))
if params.norm_id_or_name: query_params.append(("NormunNumarasiAdlar_id", params.norm_id_or_name))
if params.norm_article: query_params.append(("NormunMaddeNumarasi", params.norm_article))
if params.review_outcomes:
for outcome_val in params.review_outcomes:
if outcome_val and outcome_val != "ALL": query_params.append(("IncelemeTuruKararSonuclar_id[]", outcome_val))
if params.reason_for_final_outcome and params.reason_for_final_outcome and params.reason_for_final_outcome != "ALL":
query_params.append(("KararSonucununGerekcesi", params.reason_for_final_outcome))
if params.basis_constitution_article_numbers:
for article_no in params.basis_constitution_article_numbers: query_params.append(("DayanakHukmu[]", article_no))
if params.official_gazette_date_start: query_params.append(("ResmiGazeteTarihiIlk", params.official_gazette_date_start))
if params.official_gazette_date_end: query_params.append(("ResmiGazeteTarihiSon", params.official_gazette_date_end))
if params.official_gazette_number_start: query_params.append(("ResmiGazeteSayisiIlk", params.official_gazette_number_start))
if params.official_gazette_number_end: query_params.append(("ResmiGazeteSayisiSon", params.official_gazette_number_end))
if params.has_press_release and params.has_press_release and params.has_press_release != "ALL": query_params.append(("BasinDuyurusu", params.has_press_release))
if params.has_dissenting_opinion and params.has_dissenting_opinion and params.has_dissenting_opinion != "ALL": query_params.append(("KarsiOy", params.has_dissenting_opinion))
if params.has_different_reasoning and params.has_different_reasoning and params.has_different_reasoning != "ALL": query_params.append(("FarkliGerekce", params.has_different_reasoning))
# Add pagination and sorting parameters as query params instead of URL path
if params.results_per_page and params.results_per_page != 10:
query_params.append(("SatirSayisi", str(params.results_per_page)))
if params.sort_by_criteria and params.sort_by_criteria != "KararTarihi":
query_params.append(("Siralama", params.sort_by_criteria))
if params.page_to_fetch and params.page_to_fetch > 1:
query_params.append(("page", str(params.page_to_fetch)))
return query_params
async def search_norm_denetimi_decisions(
self,
params: AnayasaNormDenetimiSearchRequest
) -> AnayasaSearchResult:
# Use simple /Ara endpoint - the complex path structure seems to cause 404s
request_path = f"/{self.SEARCH_PATH_SEGMENT}"
final_query_params = self._build_search_query_params_for_aym(params)
logger.info(f"AnayasaMahkemesiApiClient: Performing Norm Denetimi search. Path: {request_path}, Params: {final_query_params}")
try:
response = await self.http_client.get(request_path, params=final_query_params)
response.raise_for_status()
html_content = response.text
except httpx.RequestError as e:
logger.error(f"AnayasaMahkemesiApiClient: HTTP request error during Norm Denetimi search: {e}")
raise
except Exception as e:
logger.error(f"AnayasaMahkemesiApiClient: Error processing Norm Denetimi search request: {e}")
raise
soup = BeautifulSoup(html_content, 'html.parser')
total_records = None
bulunan_karar_div = soup.find("div", class_="bulunankararsayisi")
if not bulunan_karar_div: # Fallback for mobile view
bulunan_karar_div = soup.find("div", class_="bulunankararsayisiMobil")
if bulunan_karar_div:
match_records = re.search(r'(\d+)\s*Karar Bulundu', bulunan_karar_div.get_text(strip=True))
if match_records:
total_records = int(match_records.group(1))
processed_decisions: List[AnayasaDecisionSummary] = []
decision_divs = soup.find_all("div", class_="birkarar")
for decision_div in decision_divs:
link_tag = decision_div.find("a", href=True)
doc_url_path = link_tag['href'] if link_tag else None
decision_page_url_str = urljoin(self.BASE_URL, doc_url_path) if doc_url_path else None
title_div = decision_div.find("div", class_="bkararbaslik")
ek_no_text_raw = title_div.get_text(strip=True, separator=" ").replace('\xa0', ' ') if title_div else ""
ek_no_match = re.search(r"(E\.\s*\d+/\d+\s*,\s*K\.\s*\d+/\d+)", ek_no_text_raw)
ek_no_text = ek_no_match.group(1) if ek_no_match else ek_no_text_raw.split("Sayılı Karar")[0].strip()
keyword_count_div = title_div.find("div", class_="BulunanKelimeSayisi") if title_div else None
keyword_count_text = keyword_count_div.get_text(strip=True).replace("Bulunan Kelime Sayısı", "").strip() if keyword_count_div else None
keyword_count = int(keyword_count_text) if keyword_count_text and keyword_count_text.isdigit() else None
info_div = decision_div.find("div", class_="kararbilgileri")
info_parts = [part.strip() for part in info_div.get_text(separator="|").split("|")] if info_div else []
app_type_summary = info_parts[0] if len(info_parts) > 0 else None
applicant_summary = info_parts[1] if len(info_parts) > 1 else None
outcome_summary = info_parts[2] if len(info_parts) > 2 else None
dec_date_raw = info_parts[3] if len(info_parts) > 3 else None
decision_date_summary = dec_date_raw.replace("Karar Tarihi:", "").strip() if dec_date_raw else None
reviewed_norms_list: List[AnayasaReviewedNormInfo] = []
details_table_container = decision_div.find_next_sibling("div", class_=re.compile(r"col-sm-12")) # The details table is in a sibling div
if details_table_container:
details_table = details_table_container.find("table", class_="table")
if details_table and details_table.find("tbody"):
for row in details_table.find("tbody").find_all("tr"):
cells = row.find_all("td")
if len(cells) == 6:
reviewed_norms_list.append(AnayasaReviewedNormInfo(
norm_name_or_number=cells[0].get_text(strip=True) or None,
article_number=cells[1].get_text(strip=True) or None,
review_type_and_outcome=cells[2].get_text(strip=True) or None,
outcome_reason=cells[3].get_text(strip=True) or None,
basis_constitution_articles_cited=[a.strip() for a in cells[4].get_text(strip=True).split(',') if a.strip()] if cells[4].get_text(strip=True) else [],
postponement_period=cells[5].get_text(strip=True) or None
))
processed_decisions.append(AnayasaDecisionSummary(
decision_reference_no=ek_no_text,
decision_page_url=decision_page_url_str,
keywords_found_count=keyword_count,
application_type_summary=app_type_summary,
applicant_summary=applicant_summary,
decision_outcome_summary=outcome_summary,
decision_date_summary=decision_date_summary,
reviewed_norms=reviewed_norms_list
))
return AnayasaSearchResult(
decisions=processed_decisions,
total_records_found=total_records,
retrieved_page_number=params.page_to_fetch
)
def _convert_html_to_markdown_norm_denetimi(self, full_decision_html_content: str) -> Optional[str]:
"""Converts direct HTML content from an Anayasa Mahkemesi Norm Denetimi decision page to Markdown."""
if not full_decision_html_content:
return None
processed_html = html.unescape(full_decision_html_content)
soup = BeautifulSoup(processed_html, "html.parser")
html_input_for_markdown = ""
karar_tab_content = soup.find("div", id="Karar") # "KARAR" tab content
if karar_tab_content:
karar_metni_div = karar_tab_content.find("div", class_="KararMetni")
if karar_metni_div:
# Remove scripts and styles
for script_tag in karar_metni_div.find_all("script"): script_tag.decompose()
for style_tag in karar_metni_div.find_all("style"): style_tag.decompose()
# Remove "Künye Kopyala" button and other non-content divs
for item_div in karar_metni_div.find_all("div", class_="item col-sm-12"): item_div.decompose()
for modal_div in karar_metni_div.find_all("div", class_="modal fade"): modal_div.decompose() # If any modals
word_section = karar_metni_div.find("div", class_="WordSection1")
html_input_for_markdown = str(word_section) if word_section else str(karar_metni_div)
else:
html_input_for_markdown = str(karar_tab_content)
else:
# Fallback if specific structure is not found
word_section_fallback = soup.find("div", class_="WordSection1")
if word_section_fallback:
html_input_for_markdown = str(word_section_fallback)
else:
# Last resort: use the whole body or the raw HTML
body_tag = soup.find("body")
html_input_for_markdown = str(body_tag) if body_tag else processed_html
markdown_text = None
try:
# Ensure the content is wrapped in basic HTML structure if it's not already
if not html_input_for_markdown.strip().lower().startswith(("<html", "<!doctype")):
html_content = f"<html><head><meta charset=\"UTF-8\"></head><body>{html_input_for_markdown}</body></html>"
else:
html_content = html_input_for_markdown
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_content.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_text = conversion_result.text_content
except Exception as e:
logger.error(f"AnayasaMahkemesiApiClient: MarkItDown conversion error: {e}")
return markdown_text
async def get_decision_document_as_markdown(
self,
document_url: str,
page_number: int = 1
) -> AnayasaDocumentMarkdown:
"""
Retrieves a specific Anayasa Mahkemesi (Norm Denetimi) decision,
converts its content to Markdown, and returns the requested page/chunk.
"""
full_url = urljoin(self.BASE_URL, document_url) if not document_url.startswith("http") else document_url
logger.info(f"AnayasaMahkemesiApiClient: Fetching Norm Denetimi document for Markdown (page {page_number}) from URL: {full_url}")
decision_ek_no_from_page = None
decision_date_from_page = None
official_gazette_from_page = None
try:
# Use a new client instance for document fetching if headers/timeout needs to be different,
# or reuse self.http_client if settings are compatible. For now, self.http_client.
get_response = await self.http_client.get(full_url, headers={"Accept": "text/html"})
get_response.raise_for_status()
html_content_from_api = get_response.text
if not isinstance(html_content_from_api, str) or not html_content_from_api.strip():
logger.warning(f"AnayasaMahkemesiApiClient: Received empty or non-string HTML from URL {full_url}.")
return AnayasaDocumentMarkdown(
source_url=full_url, markdown_chunk=None, current_page=page_number, total_pages=0, is_paginated=False
)
# Extract metadata from the page content (E.K. No, Date, RG)
soup = BeautifulSoup(html_content_from_api, "html.parser")
karar_metni_div = soup.find("div", class_="KararMetni") # Usually within div#Karar
if not karar_metni_div: # Fallback if not in KararMetni
karar_metni_div = soup.find("div", class_="WordSection1")
# Initialize with empty string defaults
decision_ek_no_from_page = ""
decision_date_from_page = ""
official_gazette_from_page = ""
if karar_metni_div:
# Attempt to find E.K. No (Esas No, Karar No)
# Norm Denetimi pages often have this in bold <p> tags directly or in the WordSection1
# Look for patterns like "Esas No.: YYYY/NN" and "Karar No.: YYYY/NN"
esas_no_tag = karar_metni_div.find(lambda tag: tag.name == "p" and tag.find("b") and "Esas No.:" in tag.find("b").get_text())
karar_no_tag = karar_metni_div.find(lambda tag: tag.name == "p" and tag.find("b") and "Karar No.:" in tag.find("b").get_text())
karar_tarihi_tag = karar_metni_div.find(lambda tag: tag.name == "p" and tag.find("b") and "Karar tarihi:" in tag.find("b").get_text()) # Less common on Norm pages
resmi_gazete_tag = karar_metni_div.find(lambda tag: tag.name == "p" and ("Resmî Gazete tarih ve sayısı:" in tag.get_text() or "Resmi Gazete tarih/sayı:" in tag.get_text()))
if esas_no_tag and esas_no_tag.find("b") and karar_no_tag and karar_no_tag.find("b"):
esas_str = esas_no_tag.find("b").get_text(strip=True).replace('Esas No.:', '').strip()
karar_str = karar_no_tag.find("b").get_text(strip=True).replace('Karar No.:', '').strip()
decision_ek_no_from_page = f"E.{esas_str}, K.{karar_str}"
if karar_tarihi_tag and karar_tarihi_tag.find("b"):
decision_date_from_page = karar_tarihi_tag.find("b").get_text(strip=True).replace("Karar tarihi:", "").strip()
elif karar_metni_div: # Fallback for Karar Tarihi if not in specific tag
date_match = re.search(r"Karar Tarihi\s*:\s*([\d\.]+)", karar_metni_div.get_text()) # Norm pages often use DD.MM.YYYY
if date_match: decision_date_from_page = date_match.group(1).strip()
if resmi_gazete_tag:
# Try to get the bold part first if it exists
bold_rg_tag = resmi_gazete_tag.find("b")
rg_text_content = bold_rg_tag.get_text(strip=True) if bold_rg_tag else resmi_gazete_tag.get_text(strip=True)
official_gazette_from_page = rg_text_content.replace("Resmî Gazete tarih ve sayısı:", "").replace("Resmi Gazete tarih/sayı:", "").strip()
full_markdown_content = self._convert_html_to_markdown_norm_denetimi(html_content_from_api)
if not full_markdown_content:
return AnayasaDocumentMarkdown(
source_url=full_url,
decision_reference_no_from_page=decision_ek_no_from_page,
decision_date_from_page=decision_date_from_page,
official_gazette_info_from_page=official_gazette_from_page,
markdown_chunk=None,
current_page=page_number,
total_pages=0,
is_paginated=False
)
content_length = len(full_markdown_content)
total_pages = math.ceil(content_length / self.DOCUMENT_MARKDOWN_CHUNK_SIZE)
if total_pages == 0: total_pages = 1
current_page_clamped = max(1, min(page_number, total_pages))
start_index = (current_page_clamped - 1) * self.DOCUMENT_MARKDOWN_CHUNK_SIZE
end_index = start_index + self.DOCUMENT_MARKDOWN_CHUNK_SIZE
markdown_chunk = full_markdown_content[start_index:end_index]
return AnayasaDocumentMarkdown(
source_url=full_url,
decision_reference_no_from_page=decision_ek_no_from_page,
decision_date_from_page=decision_date_from_page,
official_gazette_info_from_page=official_gazette_from_page,
markdown_chunk=markdown_chunk,
current_page=current_page_clamped,
total_pages=total_pages,
is_paginated=(total_pages > 1)
)
except httpx.RequestError as e:
logger.error(f"AnayasaMahkemesiApiClient: HTTP error fetching Norm Denetimi document from {full_url}: {e}")
raise
except Exception as e:
logger.error(f"AnayasaMahkemesiApiClient: General error processing Norm Denetimi document from {full_url}: {e}")
raise
async def close_client_session(self):
if hasattr(self, 'http_client') and self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("AnayasaMahkemesiApiClient (Norm Denetimi): HTTP client session closed.")
@@ -1,230 +0,0 @@
# anayasa_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional, Dict, Any, Literal
from enum import Enum
# --- Enums (AnayasaDonemEnum, etc. - same as before) ---
class AnayasaDonemEnum(str, Enum):
TUMU = "ALL"
DONEM_1961 = "1"
DONEM_1982 = "2"
class AnayasaVarYokEnum(str, Enum):
TUMU = "ALL"
YOK = "0"
VAR = "1"
class AnayasaIncelemeSonucuEnum(str, Enum):
TUMU = "ALL"
ESAS_ACILMAMIS_SAYILMA = "1"
ESAS_IPTAL = "2"
ESAS_KARAR_YER_OLMADIGI = "3"
ESAS_RET = "4"
ILK_ACILMAMIS_SAYILMA = "5"
ILK_ISIN_GERI_CEVRILMESI = "6"
ILK_KARAR_YER_OLMADIGI = "7"
ILK_RET = "8"
KANUN_6216_M43_4_IPTAL = "12"
class AnayasaSonucGerekcesiEnum(str, Enum):
TUMU = "ALL"
ANAYASAYA_AYKIRI_DEGIL = "29"
ANAYASAYA_ESAS_YONUNDEN_AYKIRILIK = "1"
ANAYASAYA_ESAS_YONUNDEN_UYGUNLUK = "2"
ANAYASAYA_SEKIL_ESAS_UYGUNLUK = "30"
ANAYASAYA_SEKIL_YONUNDEN_AYKIRILIK = "3"
ANAYASAYA_SEKIL_YONUNDEN_UYGUNLUK = "4"
AYKIRILIK_ANAYASAYA_ESAS_YONUNDEN_DUPLICATE = "27"
BASVURU_KARARI = "5"
DENETIM_DISI = "6"
DIGER_GEREKCE_1 = "7"
DIGER_GEREKCE_2 = "8"
EKSIKLIGIN_GIDERILMEMESI = "9"
GEREKCE = "10"
GOREV = "11"
GOREV_YETKI = "12"
GOREVLI_MAHKEME = "13"
GORULMEKTE_OLAN_DAVA = "14"
MAHKEME = "15"
NORMDA_DEGISIKLIK_YAPILMASI = "16"
NORMUN_YURURLUKTEN_KALDIRILMASI = "17"
ON_YIL_YASAGI = "18"
SURE = "19"
USULE_UYMAMA = "20"
UYGULANACAK_NORM = "21"
UYGULANAMAZ_HALE_GELME = "22"
YETKI = "23"
YETKI_SURE = "24"
YOK_HUKMUNDE_OLMAMA = "25"
YOKLUK = "26"
# --- End Enums ---
class AnayasaNormDenetimiSearchRequest(BaseModel):
"""Model for Anayasa Mahkemesi (Norm Denetimi) search request for the MCP tool."""
keywords_all: Optional[List[str]] = Field(default_factory=list, description="Keywords for AND logic (KelimeAra[]).")
keywords_any: Optional[List[str]] = Field(default_factory=list, description="Keywords for OR logic (HerhangiBirKelimeAra[]).")
keywords_exclude: Optional[List[str]] = Field(default_factory=list, description="Keywords to exclude (BulunmayanKelimeAra[]).")
period: Optional[Literal["ALL", "1", "2"]] = Field(default="ALL", description="Constitutional period (Donemler_id).")
case_number_esas: str = Field("", description="Case registry number (EsasNo), e.g., '2023/123'.")
decision_number_karar: str = Field("", description="Decision number (KararNo), e.g., '2023/456'.")
first_review_date_start: str = Field("", description="First review start date (IlkIncelemeTarihiIlk), format DD/MM/YYYY.")
first_review_date_end: str = Field("", description="First review end date (IlkIncelemeTarihiSon), format DD/MM/YYYY.")
decision_date_start: str = Field("", description="Decision start date (KararTarihiIlk), format DD/MM/YYYY.")
decision_date_end: str = Field("", description="Decision end date (KararTarihiSon), format DD/MM/YYYY.")
application_type: Optional[Literal["ALL", "1", "2", "3"]] = Field(default="ALL", description="Type of application (BasvuruTurler_id).")
applicant_general_name: str = Field("", description="General applicant name (BasvuranGeneller_id).")
applicant_specific_name: str = Field("", description="Specific applicant name (BasvuranOzeller_id).")
official_gazette_date_start: str = Field("", description="Official Gazette start date (ResmiGazeteTarihiIlk), format DD/MM/YYYY.")
official_gazette_date_end: str = Field("", description="Official Gazette end date (ResmiGazeteTarihiSon), format DD/MM/YYYY.")
official_gazette_number_start: str = Field("", description="Official Gazette starting number (ResmiGazeteSayisiIlk).")
official_gazette_number_end: str = Field("", description="Official Gazette ending number (ResmiGazeteSayisiSon).")
has_press_release: Optional[Literal["ALL", "0", "1"]] = Field(default="ALL", description="Press release available (BasinDuyurusu).")
has_dissenting_opinion: Optional[Literal["ALL", "0", "1"]] = Field(default="ALL", description="Dissenting opinion exists (KarsiOy).")
has_different_reasoning: Optional[Literal["ALL", "0", "1"]] = Field(default="ALL", description="Different reasoning exists (FarkliGerekce).")
attending_members_names: Optional[List[str]] = Field(default_factory=list, description="List of attending members' exact names (Uyeler_id[]).")
rapporteur_name: str = Field("", description="Rapporteur's exact name (Raportorler_id).")
norm_type: Optional[Literal["ALL", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "0"]] = Field(default="ALL", description="Type of the reviewed norm (NormunTurler_id).")
norm_id_or_name: str = Field("", description="Number or name of the norm (NormunNumarasiAdlar_id).")
norm_article: str = Field("", description="Article number of the norm (NormunMaddeNumarasi).")
review_outcomes: Optional[List[Literal["1", "2", "3", "4", "5", "6", "7", "8", "12"]]] = Field(default_factory=list, description="List of review types and outcomes (IncelemeTuruKararSonuclar_id[]).")
reason_for_final_outcome: Optional[Literal["ALL", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "15", "16", "17", "18", "19", "20", "21", "22", "23", "24", "25", "26", "27", "29", "30"]] = Field(default="ALL", description="Main reason for the decision outcome (KararSonucununGerekcesi).")
basis_constitution_article_numbers: Optional[List[str]] = Field(default_factory=list, description="List of supporting Constitution article numbers (DayanakHukmu[]).")
results_per_page: int = Field(10, ge=1, le=10, description="Results per page.")
page_to_fetch: int = Field(1, ge=1, description="Page number to fetch for results list.")
sort_by_criteria: str = Field("KararTarihi", description="Sort criteria. Options: 'KararTarihi', 'YayinTarihi', 'Toplam' (keyword count).")
class AnayasaReviewedNormInfo(BaseModel):
"""Details of a norm reviewed within an AYM decision summary."""
norm_name_or_number: str = Field("", description="Norm name or number")
article_number: str = Field("", description="Article number")
review_type_and_outcome: str = Field("", description="Review type and outcome")
outcome_reason: str = Field("", description="Outcome reason")
basis_constitution_articles_cited: List[str] = Field(default_factory=list)
postponement_period: str = Field("", description="Postponement period")
class AnayasaDecisionSummary(BaseModel):
"""Model for a single Anayasa Mahkemesi (Norm Denetimi) decision summary from search results."""
decision_reference_no: str = Field("", description="Decision reference number")
decision_page_url: str = Field("", description="Decision page URL")
keywords_found_count: Optional[int] = Field(0, description="Keywords found count")
application_type_summary: str = Field("", description="Application type summary")
applicant_summary: str = Field("", description="Applicant summary")
decision_outcome_summary: str = Field("", description="Decision outcome summary")
decision_date_summary: str = Field("", description="Decision date summary")
reviewed_norms: List[AnayasaReviewedNormInfo] = Field(default_factory=list)
class AnayasaSearchResult(BaseModel):
"""Model for the overall search result for Anayasa Mahkemesi Norm Denetimi decisions."""
decisions: List[AnayasaDecisionSummary]
total_records_found: int = Field(0, description="Total records found")
retrieved_page_number: int = Field(1, description="Retrieved page number")
class AnayasaDocumentMarkdown(BaseModel):
"""
Model for an Anayasa Mahkemesi (Norm Denetimi) decision document, containing a chunk of Markdown content
and pagination information.
"""
source_url: HttpUrl
decision_reference_no_from_page: str = Field("", description="E.K. No parsed from the document page.")
decision_date_from_page: str = Field("", description="Decision date parsed from the document page.")
official_gazette_info_from_page: str = Field("", description="Official Gazette info parsed from the document page.")
markdown_chunk: str = Field("", description="A 5,000 character chunk of the Markdown content.") # Corrected chunk size
current_page: int = Field(description="The current page number of the markdown chunk (1-indexed).")
total_pages: int = Field(description="Total number of pages for the full markdown content.")
is_paginated: bool = Field(description="True if the full markdown content is split into multiple pages.")
# --- Models for Anayasa Mahkemesi - Bireysel Başvuru Karar Raporu ---
class AnayasaBireyselReportSearchRequest(BaseModel):
"""Model for Anayasa Mahkemesi (Bireysel Başvuru) 'Karar Arama Raporu' search request."""
keywords: Optional[List[str]] = Field(default_factory=list, description="Keywords for AND logic (KelimeAra[]).")
page_to_fetch: int = Field(1, ge=1, description="Page number to fetch for the report (page). Default is 1.")
class AnayasaBireyselReportDecisionDetail(BaseModel):
"""Details of a specific right/claim within a Bireysel Başvuru decision summary in a report."""
hak: str = Field("", description="İhlal edildiği iddia edilen hak (örneğin, Mülkiyet hakkı).")
mudahale_iddiasi: str = Field("", description="İhlale neden olan müdahale iddiası.")
sonuc: str = Field("", description="İnceleme sonucu (örneğin, İhlal, Düşme).")
giderim: str = Field("", description="Kararlaştırılan giderim (örneğin, Yeniden yargılama).")
class AnayasaBireyselReportDecisionSummary(BaseModel):
"""Model for a single Anayasa Mahkemesi (Bireysel Başvuru) decision summary from a 'Karar Arama Raporu'."""
title: str = Field("", description="Başvurunun başlığı (e.g., 'HASAN DURMUŞ Başvurusuna İlişkin Karar').")
decision_reference_no: str = Field("", description="Başvuru Numarası (e.g., '2019/19126').")
decision_page_url: str = Field("", description="URL to the full decision page.")
decision_type_summary: str = Field("", description="Karar Türü (Başvuru Sonucu) (e.g., 'Esas (İhlal)').")
decision_making_body: str = Field("", description="Kararı Veren Birim (e.g., 'Genel Kurul', 'Birinci Bölüm').")
application_date_summary: str = Field("", description="Başvuru Tarihi (DD/MM/YYYY).")
decision_date_summary: str = Field("", description="Karar Tarihi (DD/MM/YYYY).")
application_subject_summary: str = Field("", description="Başvuru konusunun özeti.")
details: List[AnayasaBireyselReportDecisionDetail] = Field(default_factory=list, description="İncelenen haklar ve sonuçlarına ilişkin detaylar.")
class AnayasaBireyselReportSearchResult(BaseModel):
"""Model for the overall search result for Anayasa Mahkemesi 'Karar Arama Raporu'."""
decisions: List[AnayasaBireyselReportDecisionSummary]
total_records_found: int = Field(0, description="Raporda bulunan toplam karar sayısı.")
retrieved_page_number: int = Field(description="Alınan rapor sayfa numarası.")
class AnayasaBireyselBasvuruDocumentMarkdown(BaseModel):
"""
Model for an Anayasa Mahkemesi (Bireysel Başvuru) decision document, containing a chunk of Markdown content
and pagination information. Fetched from /BB/YYYY/NNNN paths.
"""
source_url: HttpUrl
basvuru_no_from_page: Optional[str] = Field(None, description="Başvuru Numarası (B.No) parsed from the document page.")
karar_tarihi_from_page: Optional[str] = Field(None, description="Decision date parsed from the document page.")
basvuru_tarihi_from_page: Optional[str] = Field(None, description="Application date parsed from the document page.")
karari_veren_birim_from_page: Optional[str] = Field(None, description="Deciding body (Bölüm/Genel Kurul) parsed from the document page.")
karar_turu_from_page: Optional[str] = Field(None, description="Decision type (Başvuru Sonucu) parsed from the document page.")
resmi_gazete_info_from_page: Optional[str] = Field(None, description="Official Gazette info parsed from the document page, if available.")
markdown_chunk: Optional[str] = Field(None, description="A 5,000 character chunk of the Markdown content.")
current_page: int = Field(description="The current page number of the markdown chunk (1-indexed).")
total_pages: int = Field(description="Total number of pages for the full markdown content.")
is_paginated: bool = Field(description="True if the full markdown content is split into multiple pages.")
# --- End Models for Bireysel Başvuru ---
# --- Unified Models ---
class AnayasaUnifiedSearchRequest(BaseModel):
"""Unified search request for both Norm Denetimi and Bireysel Başvuru."""
decision_type: Literal["norm_denetimi", "bireysel_basvuru"] = Field(..., description="Decision type: norm_denetimi or bireysel_basvuru")
# Common parameters
keywords: List[str] = Field(default_factory=list, description="Keywords to search for")
page_to_fetch: int = Field(1, ge=1, le=100, description="Page number to fetch (1-100)")
results_per_page: int = Field(10, ge=1, le=100, description="Results per page (1-100)")
# Norm Denetimi specific parameters (ignored for bireysel_basvuru)
keywords_all: List[str] = Field(default_factory=list, description="All keywords must be present (norm_denetimi only)")
keywords_any: List[str] = Field(default_factory=list, description="Any of these keywords (norm_denetimi only)")
decision_type_norm: Literal["ALL", "1", "2", "3"] = Field("ALL", description="Decision type for norm denetimi")
application_date_start: str = Field("", description="Application start date (norm_denetimi only)")
application_date_end: str = Field("", description="Application end date (norm_denetimi only)")
# Bireysel Başvuru specific parameters (ignored for norm_denetimi)
decision_start_date: str = Field("", description="Decision start date (bireysel_basvuru only)")
decision_end_date: str = Field("", description="Decision end date (bireysel_basvuru only)")
norm_type: Literal["ALL", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12", "13", "14", "0"] = Field("ALL", description="Norm type (bireysel_basvuru only)")
subject_category: str = Field("", description="Subject category (bireysel_basvuru only)")
class AnayasaUnifiedSearchResult(BaseModel):
"""Unified search result containing decisions from either system."""
decision_type: Literal["norm_denetimi", "bireysel_basvuru"] = Field(..., description="Type of decisions returned")
decisions: List[Dict[str, Any]] = Field(default_factory=list, description="Decision list (structure varies by type)")
total_records_found: int = Field(0, description="Total number of records found")
retrieved_page_number: int = Field(1, description="Page number that was retrieved")
class AnayasaUnifiedDocumentMarkdown(BaseModel):
"""Unified document model for both Norm Denetimi and Bireysel Başvuru."""
decision_type: Literal["norm_denetimi", "bireysel_basvuru"] = Field(..., description="Type of document")
source_url: HttpUrl = Field(..., description="Source URL of the document")
document_data: Dict[str, Any] = Field(default_factory=dict, description="Document content and metadata")
markdown_chunk: Optional[str] = Field(None, description="Markdown content chunk")
current_page: int = Field(1, description="Current page number")
total_pages: int = Field(1, description="Total number of pages")
is_paginated: bool = Field(False, description="Whether document is paginated")
@@ -1,122 +0,0 @@
# anayasa_mcp_module/unified_client.py
# Unified client for both Norm Denetimi and Bireysel Başvuru
import logging
from typing import Optional
from urllib.parse import urlparse
from .models import (
AnayasaUnifiedSearchRequest,
AnayasaUnifiedSearchResult,
AnayasaUnifiedDocumentMarkdown,
# Removed AnayasaDecisionTypeEnum - now using string literals
AnayasaNormDenetimiSearchRequest,
AnayasaBireyselReportSearchRequest
)
from .client import AnayasaMahkemesiApiClient
from .bireysel_client import AnayasaBireyselBasvuruApiClient
logger = logging.getLogger(__name__)
class AnayasaUnifiedClient:
"""Unified client that handles both Norm Denetimi and Bireysel Başvuru searches."""
def __init__(self, request_timeout: float = 60.0):
self.norm_client = AnayasaMahkemesiApiClient(request_timeout)
self.bireysel_client = AnayasaBireyselBasvuruApiClient(request_timeout)
async def search_unified(self, params: AnayasaUnifiedSearchRequest) -> AnayasaUnifiedSearchResult:
"""Unified search that routes to appropriate client based on decision_type."""
if params.decision_type == "norm_denetimi":
# Convert to norm denetimi request
norm_params = AnayasaNormDenetimiSearchRequest(
keywords_all=params.keywords_all or params.keywords,
keywords_any=params.keywords_any,
application_type=params.decision_type_norm,
page_to_fetch=params.page_to_fetch,
results_per_page=params.results_per_page
)
result = await self.norm_client.search_norm_denetimi_decisions(norm_params)
# Convert to unified format
decisions_list = [decision.model_dump() for decision in result.decisions]
return AnayasaUnifiedSearchResult(
decision_type="norm_denetimi",
decisions=decisions_list,
total_records_found=result.total_records_found,
retrieved_page_number=result.retrieved_page_number
)
elif params.decision_type == "bireysel_basvuru":
# Convert to bireysel başvuru request
bireysel_params = AnayasaBireyselReportSearchRequest(
keywords=params.keywords,
decision_start_date=params.decision_start_date,
decision_end_date=params.decision_end_date,
norm_type=params.norm_type,
subject_category=params.subject_category,
page_to_fetch=params.page_to_fetch,
results_per_page=params.results_per_page
)
result = await self.bireysel_client.search_bireysel_basvuru_report(bireysel_params)
# Convert to unified format
decisions_list = [decision.model_dump() for decision in result.decisions]
return AnayasaUnifiedSearchResult(
decision_type="bireysel_basvuru",
decisions=decisions_list,
total_records_found=result.total_records_found,
retrieved_page_number=result.retrieved_page_number
)
else:
raise ValueError(f"Unsupported decision type: {params.decision_type}")
async def get_document_unified(self, document_url: str, page_number: int = 1) -> AnayasaUnifiedDocumentMarkdown:
"""Unified document retrieval that auto-detects the appropriate client."""
# Auto-detect decision type based on URL
parsed_url = urlparse(document_url)
if "normkararlarbilgibankasi" in parsed_url.netloc or "/ND/" in document_url:
# Norm Denetimi document
result = await self.norm_client.get_decision_document_as_markdown(document_url, page_number)
return AnayasaUnifiedDocumentMarkdown(
decision_type="norm_denetimi",
source_url=result.source_url,
document_data=result.model_dump(),
markdown_chunk=result.markdown_chunk,
current_page=result.current_page,
total_pages=result.total_pages,
is_paginated=result.is_paginated
)
elif "kararlarbilgibankasi" in parsed_url.netloc or "/BB/" in document_url:
# Bireysel Başvuru document
result = await self.bireysel_client.get_decision_document_as_markdown(document_url, page_number)
return AnayasaUnifiedDocumentMarkdown(
decision_type="bireysel_basvuru",
source_url=result.source_url,
document_data=result.model_dump(),
markdown_chunk=result.markdown_chunk,
current_page=result.current_page,
total_pages=result.total_pages,
is_paginated=result.is_paginated
)
else:
raise ValueError(f"Cannot determine document type from URL: {document_url}")
async def close_client_session(self):
"""Close both client sessions."""
if hasattr(self.norm_client, 'close_client_session'):
await self.norm_client.close_client_session()
if hasattr(self.bireysel_client, 'close_client_session'):
await self.bireysel_client.close_client_session()
-566
View File
@@ -1,566 +0,0 @@
"""
ASGI application for Yargı MCP Server
This module provides ASGI/HTTP access to the Yargı MCP server,
allowing it to be deployed as a web service with FastAPI wrapper
for Stripe webhook integration.
Usage:
uvicorn asgi_app:app --host 0.0.0.0 --port 8000
"""
import os
import time
import logging
from datetime import datetime, timedelta
from fastapi import FastAPI, Request, HTTPException, Query
from fastapi.responses import JSONResponse, HTMLResponse
from fastapi.exception_handlers import http_exception_handler
from starlette.middleware import Middleware
from starlette.middleware.cors import CORSMiddleware
from starlette.responses import Response
from starlette.requests import Request as StarletteRequest
# Import the MCP app creator function
from mcp_server_main import create_app
# Import Stripe webhook router
from stripe_webhook import router as stripe_router
# Import simplified MCP Auth HTTP adapter
from mcp_auth_http_simple import router as mcp_auth_router
# OAuth configuration from environment variables
CLERK_ISSUER = os.getenv("CLERK_ISSUER", "https://accounts.yargimcp.com")
BASE_URL = os.getenv("BASE_URL", "https://yargimcp.com")
# Setup logging
logger = logging.getLogger(__name__)
# Configure CORS and Auth middleware
cors_origins = os.getenv("ALLOWED_ORIGINS", "*").split(",")
# Import FastMCP Bearer Auth Provider
from fastmcp.server.auth import BearerAuthProvider
from fastmcp.server.auth.providers.bearer import RSAKeyPair
# Clerk JWT configuration for Bearer token validation
CLERK_SECRET_KEY = os.getenv("CLERK_SECRET_KEY")
CLERK_ISSUER = os.getenv("CLERK_ISSUER", "https://accounts.yargimcp.com")
CLERK_PUBLISHABLE_KEY = os.getenv("CLERK_PUBLISHABLE_KEY")
# Configure Bearer token authentication
bearer_auth = None
if CLERK_SECRET_KEY and CLERK_ISSUER:
# Production: Use Clerk JWKS endpoint for token validation
bearer_auth = BearerAuthProvider(
jwks_uri=f"{CLERK_ISSUER}/.well-known/jwks.json",
issuer=CLERK_ISSUER,
algorithm="RS256",
audience=None, # Disable audience validation - Clerk uses different audience format
required_scopes=[] # Disable scope validation - Clerk JWT has ['read', 'search']
)
logger.info(f"Bearer auth configured with Clerk JWKS: {CLERK_ISSUER}/.well-known/jwks.json")
else:
# Development: Generate RSA key pair for testing
logger.warning("No Clerk credentials found - using development RSA key pair")
dev_key_pair = RSAKeyPair.generate()
bearer_auth = BearerAuthProvider(
public_key=dev_key_pair.public_key,
issuer="https://dev.yargimcp.com",
audience="dev-mcp-server",
required_scopes=["yargi.read"]
)
# Generate a test token for development
dev_token = dev_key_pair.create_token(
subject="dev-user",
issuer="https://dev.yargimcp.com",
audience="dev-mcp-server",
scopes=["yargi.read", "yargi.search"],
expires_in_seconds=3600 * 24 # 24 hours for development
)
logger.info(f"Development Bearer token: {dev_token}")
custom_middleware = [
Middleware(
CORSMiddleware,
allow_origins=cors_origins,
allow_credentials=True,
allow_methods=["GET", "POST", "OPTIONS", "DELETE"],
allow_headers=["Content-Type", "Authorization", "X-Request-ID", "X-Session-ID"],
),
]
# Create MCP app with Bearer authentication
mcp_server = create_app(auth=bearer_auth)
# Add Starlette middleware to FastAPI (not MCP)
# MCP already has Bearer auth, no need for additional middleware on MCP level
# Create MCP Starlette sub-application with root path - mount will add /mcp prefix
mcp_app = mcp_server.http_app(path="/")
# Configure JSON encoder for proper Turkish character support
import json
from fastapi.responses import JSONResponse
class UTF8JSONResponse(JSONResponse):
def __init__(self, content=None, status_code=200, headers=None, **kwargs):
if headers is None:
headers = {}
headers["Content-Type"] = "application/json; charset=utf-8"
super().__init__(content, status_code, headers, **kwargs)
def render(self, content) -> bytes:
return json.dumps(
content,
ensure_ascii=False,
allow_nan=False,
indent=None,
separators=(",", ":"),
).encode("utf-8")
# Create FastAPI wrapper application
app = FastAPI(
title="Yargı MCP Server",
description="MCP server for Turkish legal databases with OAuth authentication",
version="0.1.0",
middleware=custom_middleware,
default_response_class=UTF8JSONResponse # Use UTF-8 JSON encoder
)
# Add Stripe webhook router to FastAPI
app.include_router(stripe_router, prefix="/api")
# Add MCP Auth HTTP adapter to FastAPI (handles OAuth endpoints)
app.include_router(mcp_auth_router)
# Custom 401 exception handler for MCP spec compliance
@app.exception_handler(401)
async def custom_401_handler(request: Request, exc: HTTPException):
"""Custom 401 handler that adds WWW-Authenticate header as required by MCP spec"""
response = await http_exception_handler(request, exc)
# Add WWW-Authenticate header pointing to protected resource metadata
# as required by RFC 9728 Section 5.1 and MCP Authorization spec
response.headers["WWW-Authenticate"] = (
'Bearer '
'error="invalid_token", '
'error_description="The access token is missing or invalid", '
f'resource="{BASE_URL}/.well-known/oauth-protected-resource"'
)
return response
# FastAPI health check endpoint - BEFORE mounting MCP app
@app.get("/health")
async def health_check():
"""Health check endpoint for monitoring"""
return JSONResponse({
"status": "healthy",
"service": "Yargı MCP Server",
"version": "0.1.0",
"tools_count": len(mcp_server._tool_manager._tools),
"auth_enabled": os.getenv("ENABLE_AUTH", "false").lower() == "true"
})
# Add explicit redirect for /mcp to /mcp/ with method preservation
@app.api_route("/mcp", methods=["GET", "POST", "HEAD", "OPTIONS"])
async def redirect_to_slash(request: Request):
"""Redirect /mcp to /mcp/ preserving HTTP method with 308"""
from fastapi.responses import RedirectResponse
return RedirectResponse(url="/mcp/", status_code=308)
# Mount MCP app at /mcp/ with trailing slash
app.mount("/mcp/", mcp_app)
# Set the lifespan context after mounting
app.router.lifespan_context = mcp_app.lifespan
# SSE transport deprecated - removed
# FastAPI root endpoint
@app.get("/")
async def root():
"""Root endpoint with service information"""
return JSONResponse({
"service": "Yargı MCP Server",
"description": "MCP server for Turkish legal databases with OAuth authentication",
"endpoints": {
"mcp": "/mcp",
"health": "/health",
"status": "/status",
"stripe_webhook": "/api/stripe/webhook",
"oauth_login": "/auth/login",
"oauth_callback": "/auth/callback",
"oauth_google": "/auth/google/login",
"user_info": "/auth/user"
},
"transports": {
"http": "/mcp"
},
"supported_databases": [
"Yargıtay (Court of Cassation)",
"Danıştay (Council of State)",
"Emsal (Precedent)",
"Uyuşmazlık Mahkemesi (Court of Jurisdictional Disputes)",
"Anayasa Mahkemesi (Constitutional Court)",
"Kamu İhale Kurulu (Public Procurement Authority)",
"Rekabet Kurumu (Competition Authority)",
"Sayıştay (Court of Accounts)",
"Bedesten API (Multiple courts)"
],
"authentication": {
"enabled": os.getenv("ENABLE_AUTH", "false").lower() == "true",
"type": "OAuth 2.0 via Clerk",
"issuer": os.getenv("CLERK_ISSUER", "https://clerk.accounts.dev"),
"providers": ["google"],
"flow": "authorization_code"
}
})
# OAuth 2.0 Authorization Server Metadata proxy (for MCP clients that can't reach Clerk directly)
# MCP Auth Toolkit expects this to be under /mcp/.well-known/oauth-authorization-server
@app.get("/mcp/.well-known/oauth-authorization-server")
async def oauth_authorization_server():
"""OAuth 2.0 Authorization Server Metadata proxy to Clerk - MCP Auth Toolkit standard location"""
return JSONResponse({
"issuer": BASE_URL,
"authorization_endpoint": "https://yargimcp.com/mcp-callback",
"token_endpoint": f"{BASE_URL}/token",
"jwks_uri": f"{CLERK_ISSUER}/.well-known/jwks.json",
"response_types_supported": ["code"],
"grant_types_supported": ["authorization_code", "refresh_token"],
"token_endpoint_auth_methods_supported": ["client_secret_basic", "none"],
"scopes_supported": ["read", "search", "openid", "profile", "email"],
"subject_types_supported": ["public"],
"id_token_signing_alg_values_supported": ["RS256"],
"claims_supported": ["sub", "iss", "aud", "exp", "iat", "email", "name"],
"code_challenge_methods_supported": ["S256"],
"service_documentation": f"{BASE_URL}/mcp",
"registration_endpoint": f"{BASE_URL}/register",
"resource_documentation": f"{BASE_URL}/mcp"
})
# Claude AI MCP specific endpoint format
@app.get("/.well-known/oauth-authorization-server/mcp")
async def oauth_authorization_server_mcp_suffix():
"""OAuth 2.0 Authorization Server Metadata - Claude AI MCP specific format"""
return JSONResponse({
"issuer": BASE_URL,
"authorization_endpoint": "https://yargimcp.com/mcp-callback",
"token_endpoint": f"{BASE_URL}/token",
"jwks_uri": f"{CLERK_ISSUER}/.well-known/jwks.json",
"response_types_supported": ["code"],
"grant_types_supported": ["authorization_code", "refresh_token"],
"token_endpoint_auth_methods_supported": ["client_secret_basic", "none"],
"scopes_supported": ["read", "search", "openid", "profile", "email"],
"subject_types_supported": ["public"],
"id_token_signing_alg_values_supported": ["RS256"],
"claims_supported": ["sub", "iss", "aud", "exp", "iat", "email", "name"],
"code_challenge_methods_supported": ["S256"],
"service_documentation": f"{BASE_URL}/mcp",
"registration_endpoint": f"{BASE_URL}/register",
"resource_documentation": f"{BASE_URL}/mcp"
})
@app.get("/.well-known/oauth-protected-resource/mcp")
async def oauth_protected_resource_mcp_suffix():
"""OAuth 2.0 Protected Resource Metadata - Claude AI MCP specific format"""
return JSONResponse({
"resource": BASE_URL,
"authorization_servers": [
BASE_URL
],
"scopes_supported": ["read", "search"],
"bearer_methods_supported": ["header"],
"resource_documentation": f"{BASE_URL}/mcp",
"resource_policy_uri": f"{BASE_URL}/privacy"
})
# Keep root level for compatibility with some MCP clients
@app.get("/.well-known/oauth-authorization-server")
async def oauth_authorization_server_root():
"""OAuth 2.0 Authorization Server Metadata proxy to Clerk - root level for compatibility"""
return JSONResponse({
"issuer": BASE_URL,
"authorization_endpoint": "https://yargimcp.com/mcp-callback",
"token_endpoint": f"{BASE_URL}/token",
"jwks_uri": f"{CLERK_ISSUER}/.well-known/jwks.json",
"response_types_supported": ["code"],
"grant_types_supported": ["authorization_code", "refresh_token"],
"token_endpoint_auth_methods_supported": ["client_secret_basic", "none"],
"scopes_supported": ["read", "search", "openid", "profile", "email"],
"subject_types_supported": ["public"],
"id_token_signing_alg_values_supported": ["RS256"],
"claims_supported": ["sub", "iss", "aud", "exp", "iat", "email", "name"],
"code_challenge_methods_supported": ["S256"],
"service_documentation": f"{BASE_URL}/mcp",
"registration_endpoint": f"{BASE_URL}/register",
"resource_documentation": f"{BASE_URL}/mcp"
})
# Note: GET /mcp is handled by the mounted MCP app itself
# This prevents 405 Method Not Allowed errors on POST requests
# OAuth 2.0 Protected Resource Metadata (RFC 9728) - MCP Spec Required
@app.get("/.well-known/oauth-protected-resource")
async def oauth_protected_resource():
"""OAuth 2.0 Protected Resource Metadata as required by MCP spec"""
return JSONResponse({
"resource": BASE_URL,
"authorization_servers": [
BASE_URL
],
"scopes_supported": ["read", "search"],
"bearer_methods_supported": ["header"],
"resource_documentation": f"{BASE_URL}/mcp",
"resource_policy_uri": f"{BASE_URL}/privacy"
})
# Standard well-known discovery endpoint
@app.get("/.well-known/mcp")
async def well_known_mcp():
"""Standard MCP discovery endpoint"""
return JSONResponse({
"mcp_server": {
"name": "Yargı MCP Server",
"version": "0.1.0",
"endpoint": f"{BASE_URL}/mcp",
"authentication": {
"type": "oauth2",
"authorization_url": f"{BASE_URL}/auth/login",
"scopes": ["read", "search"]
},
"capabilities": ["tools", "resources"],
"tools_count": len(mcp_server._tool_manager._tools)
}
})
# MCP Discovery endpoint for ChatGPT integration
@app.get("/mcp/discovery")
async def mcp_discovery():
"""MCP Discovery endpoint for ChatGPT and other MCP clients"""
return JSONResponse({
"name": "Yargı MCP Server",
"description": "MCP server for Turkish legal databases",
"version": "0.1.0",
"protocol": "mcp",
"transport": "http",
"endpoint": "/mcp",
"authentication": {
"type": "oauth2",
"authorization_url": "/auth/login",
"token_url": "/auth/callback",
"scopes": ["read", "search"],
"provider": "clerk"
},
"capabilities": {
"tools": True,
"resources": True,
"prompts": False
},
"tools_count": len(mcp_server._tool_manager._tools),
"contact": {
"url": BASE_URL,
"email": "support@yargi-mcp.dev"
}
})
# FastAPI status endpoint
@app.get("/status")
async def status():
"""Status endpoint with detailed information"""
tools = []
for tool in mcp_server._tool_manager._tools.values():
tools.append({
"name": tool.name,
"description": tool.description[:100] + "..." if len(tool.description) > 100 else tool.description
})
return JSONResponse({
"status": "operational",
"tools": tools,
"total_tools": len(tools),
"transport": "streamable_http",
"architecture": "FastAPI wrapper + MCP Starlette sub-app",
"auth_status": "enabled" if os.getenv("ENABLE_AUTH", "false").lower() == "true" else "disabled"
})
# Note: JWT token validation is now handled entirely by Clerk
# All authentication flows use Clerk JWT tokens directly
async def validate_clerk_session(request: Request, clerk_token: str = None) -> str:
"""Validate Clerk session from cookies or JWT token and return user_id"""
logger.info(f"Validating Clerk session - token provided: {bool(clerk_token)}")
try:
# Try to import Clerk SDK
from clerk_backend_api import Clerk
clerk = Clerk(bearer_auth=os.getenv("CLERK_SECRET_KEY"))
# Try JWT token first (from URL parameter)
if clerk_token:
logger.info("Validating Clerk JWT token from URL parameter")
try:
# Extract session_id from JWT token and verify with Clerk
import jwt
decoded_token = jwt.decode(clerk_token, options={"verify_signature": False})
session_id = decoded_token.get("sid") # Use standard JWT 'sid' claim
if session_id:
# Verify with Clerk using session_id
session = clerk.sessions.verify(session_id=session_id, token=clerk_token)
user_id = session.user_id if session else None
if user_id:
logger.info(f"JWT token validation successful - user_id: {user_id}")
return user_id
else:
logger.error("JWT token validation failed - no user_id in session")
else:
logger.error("No session_id found in JWT token")
except Exception as e:
logger.error(f"JWT token validation failed: {str(e)}")
# Fall through to cookie validation
# Fallback to cookie validation
logger.info("Attempting cookie-based session validation")
clerk_session = request.cookies.get("__session")
if not clerk_session:
logger.error("No Clerk session cookie found")
raise HTTPException(status_code=401, detail="No Clerk session found")
# Validate session with Clerk
session = clerk.sessions.verify_session(clerk_session)
logger.info(f"Cookie session validation successful - user_id: {session.user_id}")
return session.user_id
except ImportError:
# Fallback for development without Clerk SDK
logger.warning("Clerk SDK not available - using development fallback")
return "dev_user_123"
except Exception as e:
logger.error(f"Session validation failed: {str(e)}")
raise HTTPException(status_code=401, detail=f"Session validation failed: {str(e)}")
# MCP OAuth Callback Endpoint
@app.get("/auth/mcp-callback")
async def mcp_oauth_callback(request: Request, clerk_token: str = Query(None)):
"""Handle OAuth callback for MCP token generation"""
logger.info(f"MCP OAuth callback - clerk_token provided: {bool(clerk_token)}")
try:
# Validate Clerk session with JWT token support
user_id = await validate_clerk_session(request, clerk_token)
logger.info(f"User authenticated successfully - user_id: {user_id}")
# Use the Clerk JWT token directly (no need to generate custom token)
logger.info("User authenticated successfully via Clerk")
# Return success response
return HTMLResponse(f"""
<html>
<head>
<title>MCP Connection Successful</title>
<style>
body {{ font-family: Arial, sans-serif; text-align: center; padding: 50px; }}
.success {{ color: #28a745; }}
.token {{ background: #f8f9fa; padding: 15px; border-radius: 5px; margin: 20px 0; word-break: break-all; }}
</style>
</head>
<body>
<h1 class="success"> MCP Connection Successful!</h1>
<p>Your Yargı MCP integration is now active.</p>
<div class="token">
<strong>Authentication:</strong><br>
<code>Use your Clerk JWT token directly with Bearer authentication</code>
</div>
<p>You can now close this window and return to your MCP client.</p>
<script>
// Try to close the popup if opened as such
if (window.opener) {{
window.opener.postMessage({{
type: 'MCP_AUTH_SUCCESS',
token: 'use_clerk_jwt_token'
}}, '*');
setTimeout(() => window.close(), 3000);
}}
</script>
</body>
</html>
""")
except HTTPException as e:
logger.error(f"MCP OAuth callback failed: {e.detail}")
return HTMLResponse(f"""
<html>
<head>
<title>MCP Connection Failed</title>
<style>
body {{ font-family: Arial, sans-serif; text-align: center; padding: 50px; }}
.error {{ color: #dc3545; }}
.debug {{ background: #f8f9fa; padding: 10px; margin: 20px 0; border-radius: 5px; font-family: monospace; }}
</style>
</head>
<body>
<h1 class="error"> MCP Connection Failed</h1>
<p>{e.detail}</p>
<div class="debug">
<strong>Debug Info:</strong><br>
Clerk Token: {'✅ Provided' if clerk_token else '❌ Missing'}<br>
Error: {e.detail}<br>
Status: {e.status_code}
</div>
<p>Please try again or contact support.</p>
<a href="https://yargimcp.com/sign-in">Return to Sign In</a>
</body>
</html>
""", status_code=e.status_code)
except Exception as e:
logger.error(f"Unexpected error in MCP OAuth callback: {str(e)}")
return HTMLResponse(f"""
<html>
<head>
<title>MCP Connection Error</title>
<style>
body {{ font-family: Arial, sans-serif; text-align: center; padding: 50px; }}
.error {{ color: #dc3545; }}
</style>
</head>
<body>
<h1 class="error"> Unexpected Error</h1>
<p>An unexpected error occurred during authentication.</p>
<p>Error: {str(e)}</p>
<a href="https://yargimcp.com/sign-in">Return to Sign In</a>
</body>
</html>
""", status_code=500)
# OAuth2 Token Endpoint - Now uses Clerk JWT tokens directly
@app.post("/auth/mcp-token")
async def mcp_token_endpoint(request: Request):
"""OAuth2 token endpoint for MCP clients - returns Clerk JWT token info"""
try:
# Validate Clerk session
user_id = await validate_clerk_session(request)
return JSONResponse({
"message": "Use your Clerk JWT token directly with Bearer authentication",
"token_type": "Bearer",
"scope": "yargi.read",
"user_id": user_id,
"instructions": "Include 'Authorization: Bearer YOUR_CLERK_JWT_TOKEN' in your requests"
})
except HTTPException as e:
return JSONResponse(
status_code=e.status_code,
content={"error": "invalid_request", "error_description": e.detail}
)
# Note: Only HTTP transport supported - SSE transport deprecated
# Export for uvicorn
__all__ = ["app"]
@@ -1,17 +0,0 @@
# bddk_mcp_module/__init__.py
from .client import BddkApiClient
from .models import (
BddkSearchRequest,
BddkDecisionSummary,
BddkSearchResult,
BddkDocumentMarkdown
)
__all__ = [
"BddkApiClient",
"BddkSearchRequest",
"BddkDecisionSummary",
"BddkSearchResult",
"BddkDocumentMarkdown"
]
@@ -1,247 +0,0 @@
# bddk_mcp_module/client.py
import httpx
from typing import List, Optional, Dict, Any
import logging
import os
import re
import io
import math
from urllib.parse import urlparse
from markitdown import MarkItDown
from .models import (
BddkSearchRequest,
BddkDecisionSummary,
BddkSearchResult,
BddkDocumentMarkdown
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
class BddkApiClient:
"""
API client for searching and retrieving BDDK (Banking Regulation Authority) decisions
using Tavily Search API for discovery and direct HTTP requests for content retrieval.
"""
TAVILY_API_URL = "https://api.tavily.com/search"
BDDK_BASE_URL = "https://www.bddk.org.tr"
DOCUMENT_URL_TEMPLATE = "https://www.bddk.org.tr/Mevzuat/DokumanGetir/{document_id}"
DOCUMENT_MARKDOWN_CHUNK_SIZE = 5000 # Character limit per page
def __init__(self, request_timeout: float = 60.0):
"""Initialize the BDDK API client."""
self.tavily_api_key = os.getenv("TAVILY_API_KEY")
if not self.tavily_api_key:
# Fallback to development token
self.tavily_api_key = "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"
logger.info("Using fallback Tavily API token (development token)")
else:
logger.info("Using Tavily API key from environment variable")
self.http_client = httpx.AsyncClient(
headers={
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36"
},
timeout=httpx.Timeout(request_timeout)
)
self.markitdown = MarkItDown()
async def close_client_session(self):
"""Close the HTTP client session."""
await self.http_client.aclose()
logger.info("BddkApiClient: HTTP client session closed.")
def _extract_document_id(self, url: str) -> Optional[str]:
"""Extract document ID from BDDK URL."""
# Primary pattern: https://www.bddk.org.tr/Mevzuat/DokumanGetir/310
match = re.search(r'/DokumanGetir/(\d+)', url)
if match:
return match.group(1)
# Alternative patterns for different BDDK URL formats
# Pattern: /Liste/55 -> use as document ID
match = re.search(r'/Liste/(\d+)', url)
if match:
return match.group(1)
# Pattern: /EkGetir/13?ekId=381 -> use ekId as document ID
match = re.search(r'ekId=(\d+)', url)
if match:
return match.group(1)
return None
async def search_decisions(
self,
request: BddkSearchRequest
) -> BddkSearchResult:
"""
Search for BDDK decisions using Tavily API.
Args:
request: Search request parameters
Returns:
BddkSearchResult with matching decisions
"""
try:
headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {self.tavily_api_key}"
}
# Tavily API request - enhanced for BDDK decision documents
query = f"{request.keywords} \"Karar Sayısı\""
payload = {
"query": query,
"country": "turkey",
"include_domains": ["https://www.bddk.org.tr/Mevzuat/DokumanGetir"],
"max_results": request.pageSize,
"search_depth": "advanced"
}
# Calculate offset for pagination
if request.page > 1:
# Tavily doesn't have direct pagination, so we'll need to handle this
# For now, we'll just return empty for pages > 1
logger.warning(f"Tavily API doesn't support pagination. Page {request.page} requested.")
response = await self.http_client.post(
self.TAVILY_API_URL,
json=payload,
headers=headers
)
response.raise_for_status()
data = response.json()
# Log raw Tavily response for debugging
logger.info(f"Tavily returned {len(data.get('results', []))} results")
# Convert Tavily results to our format
decisions = []
for result in data.get("results", []):
# Extract document ID from URL
url = result.get("url", "")
logger.debug(f"Processing URL: {url}")
doc_id = self._extract_document_id(url)
if doc_id:
decision = BddkDecisionSummary(
title=result.get("title", "").replace("[PDF] ", "").strip(),
document_id=doc_id,
content=result.get("content", "")[:500] # Limit content length
)
decisions.append(decision)
logger.debug(f"Added decision: {decision.title} (ID: {doc_id})")
else:
logger.warning(f"Could not extract document ID from URL: {url}")
return BddkSearchResult(
decisions=decisions,
total_results=len(data.get("results", [])),
page=request.page,
pageSize=request.pageSize
)
except httpx.HTTPStatusError as e:
logger.error(f"HTTP error searching BDDK decisions: {e}")
if e.response.status_code == 401:
raise Exception("Tavily API authentication failed. Check API key.")
raise Exception(f"Failed to search BDDK decisions: {str(e)}")
except Exception as e:
logger.error(f"Error searching BDDK decisions: {e}")
raise Exception(f"Failed to search BDDK decisions: {str(e)}")
async def get_document_markdown(
self,
document_id: str,
page_number: int = 1
) -> BddkDocumentMarkdown:
"""
Retrieve a BDDK document and convert it to Markdown format.
Args:
document_id: BDDK document ID (e.g., '310')
page_number: Page number for paginated content (1-indexed)
Returns:
BddkDocumentMarkdown with paginated content
"""
try:
# Try different URL patterns for BDDK documents
potential_urls = [
f"https://www.bddk.org.tr/Mevzuat/DokumanGetir/{document_id}",
f"https://www.bddk.org.tr/Mevzuat/Liste/{document_id}",
f"https://www.bddk.org.tr/KurumHakkinda/EkGetir/13?ekId={document_id}",
f"https://www.bddk.org.tr/KurumHakkinda/EkGetir/5?ekId={document_id}"
]
document_url = None
response = None
# Try each URL pattern until one works
for url in potential_urls:
try:
logger.info(f"Trying BDDK document URL: {url}")
response = await self.http_client.get(
url,
follow_redirects=True
)
response.raise_for_status()
document_url = url
break
except httpx.HTTPStatusError:
continue
if not response or not document_url:
raise Exception(f"Could not find document with ID {document_id}")
logger.info(f"Successfully fetched BDDK document from: {document_url}")
# Determine content type
content_type = response.headers.get("content-type", "").lower()
# Convert to Markdown based on content type
if "pdf" in content_type:
# Handle PDF documents
pdf_stream = io.BytesIO(response.content)
result = self.markitdown.convert_stream(pdf_stream, file_extension=".pdf")
markdown_content = result.text_content
else:
# Handle HTML documents
html_stream = io.BytesIO(response.content)
result = self.markitdown.convert_stream(html_stream, file_extension=".html")
markdown_content = result.text_content
# Clean up the markdown content
markdown_content = markdown_content.strip()
# Calculate pagination
total_length = len(markdown_content)
total_pages = math.ceil(total_length / self.DOCUMENT_MARKDOWN_CHUNK_SIZE)
# Extract the requested page
start_idx = (page_number - 1) * self.DOCUMENT_MARKDOWN_CHUNK_SIZE
end_idx = start_idx + self.DOCUMENT_MARKDOWN_CHUNK_SIZE
page_content = markdown_content[start_idx:end_idx]
return BddkDocumentMarkdown(
document_id=document_id,
markdown_content=page_content,
page_number=page_number,
total_pages=total_pages
)
except httpx.HTTPStatusError as e:
logger.error(f"HTTP error fetching BDDK document {document_id}: {e}")
raise Exception(f"Failed to fetch BDDK document: {str(e)}")
except Exception as e:
logger.error(f"Error processing BDDK document {document_id}: {e}")
raise Exception(f"Failed to process BDDK document: {str(e)}")
@@ -1,43 +0,0 @@
# bddk_mcp_module/models.py
from pydantic import BaseModel, Field
from typing import List, Optional
class BddkSearchRequest(BaseModel):
"""
Request model for searching BDDK decisions via Tavily API.
BDDK (Bankacılık Düzenleme ve Denetleme Kurumu) is Turkey's Banking
Regulation and Supervision Agency responsible for banking licenses,
electronic money institutions, and financial regulations.
"""
keywords: str = Field(..., description="Search keywords in Turkish")
page: int = Field(1, ge=1, description="Page number (1-indexed)")
pageSize: int = Field(10, ge=1, le=50, description="Results per page (1-50)")
class BddkDecisionSummary(BaseModel):
"""Summary of a BDDK decision from search results."""
title: str = Field(..., description="Decision title")
document_id: str = Field(..., description="BDDK document ID (e.g., '310')")
content: str = Field(..., description="Decision summary/excerpt")
class BddkSearchResult(BaseModel):
"""Response model for BDDK decision search results."""
decisions: List[BddkDecisionSummary] = Field(
default_factory=list,
description="List of matching BDDK decisions"
)
total_results: int = Field(0, description="Total number of results")
page: int = Field(1, description="Current page number")
pageSize: int = Field(10, description="Results per page")
class BddkDocumentMarkdown(BaseModel):
"""
BDDK decision document converted to Markdown format.
Supports paginated content for long documents (5000 chars per page).
"""
document_id: str = Field(..., description="BDDK document ID")
markdown_content: str = Field("", description="Document content in Markdown")
page_number: int = Field(1, description="Current page number")
total_pages: int = Field(1, description="Total number of pages")
@@ -1 +0,0 @@
# bedesten_mcp_module/__init__.py
@@ -1,181 +0,0 @@
# bedesten_mcp_module/client.py
import httpx
import base64
from typing import Optional
import logging
from markitdown import MarkItDown
import io
from .models import (
BedestenSearchRequest, BedestenSearchResponse,
BedestenDocumentRequest, BedestenDocumentResponse,
BedestenDocumentMarkdown, BedestenDocumentRequestData
)
from .enums import get_full_birim_adi
logger = logging.getLogger(__name__)
class BedestenApiClient:
"""
API Client for Bedesten (bedesten.adalet.gov.tr) - Alternative legal decision search system.
Currently used for Yargıtay decisions, but can be extended for other court types.
"""
BASE_URL = "https://bedesten.adalet.gov.tr"
SEARCH_ENDPOINT = "/emsal-karar/searchDocuments"
DOCUMENT_ENDPOINT = "/emsal-karar/getDocumentContent"
def __init__(self, request_timeout: float = 60.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Accept": "*/*",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"AdaletApplicationName": "UyapMevzuat",
"Content-Type": "application/json; charset=utf-8",
"Origin": "https://mevzuat.adalet.gov.tr",
"Referer": "https://mevzuat.adalet.gov.tr/",
"Sec-Fetch-Dest": "empty",
"Sec-Fetch-Mode": "cors",
"Sec-Fetch-Site": "same-site",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36"
},
timeout=request_timeout
)
async def search_documents(self, search_request: BedestenSearchRequest) -> BedestenSearchResponse:
"""
Search for documents using Bedesten API.
Currently supports: YARGITAYKARARI, DANISTAYKARARI, YERELHUKMAHKARARI, etc.
"""
logger.info(f"BedestenApiClient: Searching documents with phrase: {search_request.data.phrase}")
# Map abbreviated birimAdi to full Turkish name before sending to API
original_birim_adi = search_request.data.birimAdi
mapped_birim_adi = get_full_birim_adi(original_birim_adi)
search_request.data.birimAdi = mapped_birim_adi
if original_birim_adi != "ALL":
logger.info(f"BedestenApiClient: Mapped birimAdi '{original_birim_adi}' to '{mapped_birim_adi}'")
try:
# Create request dict and remove birimAdi if empty
request_dict = search_request.model_dump()
if not request_dict["data"]["birimAdi"]: # Remove if empty string
del request_dict["data"]["birimAdi"]
response = await self.http_client.post(
self.SEARCH_ENDPOINT,
json=request_dict
)
response.raise_for_status()
response_json = response.json()
# Parse and return the response
return BedestenSearchResponse(**response_json)
except httpx.RequestError as e:
logger.error(f"BedestenApiClient: HTTP request error during search: {e}")
raise
except Exception as e:
logger.error(f"BedestenApiClient: Error processing search response: {e}")
raise
async def get_document_as_markdown(self, document_id: str) -> BedestenDocumentMarkdown:
"""
Get document content and convert to markdown.
Handles both HTML (text/html) and PDF (application/pdf) content types.
"""
logger.info(f"BedestenApiClient: Fetching document for markdown conversion (ID: {document_id})")
try:
# Prepare request
doc_request = BedestenDocumentRequest(
data=BedestenDocumentRequestData(documentId=document_id)
)
# Get document
response = await self.http_client.post(
self.DOCUMENT_ENDPOINT,
json=doc_request.model_dump()
)
response.raise_for_status()
response_json = response.json()
doc_response = BedestenDocumentResponse(**response_json)
# Decode base64 content
content_bytes = base64.b64decode(doc_response.data.content)
mime_type = doc_response.data.mimeType
logger.info(f"BedestenApiClient: Document mime type: {mime_type}")
# Convert to markdown based on mime type
if mime_type == "text/html":
html_content = content_bytes.decode('utf-8')
markdown_content = self._convert_html_to_markdown(html_content)
elif mime_type == "application/pdf":
markdown_content = self._convert_pdf_to_markdown(content_bytes)
else:
logger.warning(f"Unsupported mime type: {mime_type}")
markdown_content = f"Unsupported content type: {mime_type}. Unable to convert to markdown."
return BedestenDocumentMarkdown(
documentId=document_id,
markdown_content=markdown_content,
source_url=f"{self.BASE_URL}/document/{document_id}",
mime_type=mime_type
)
except httpx.RequestError as e:
logger.error(f"BedestenApiClient: HTTP error fetching document {document_id}: {e}")
raise
except Exception as e:
logger.error(f"BedestenApiClient: Error processing document {document_id}: {e}")
raise
def _convert_html_to_markdown(self, html_content: str) -> Optional[str]:
"""Convert HTML to Markdown using MarkItDown"""
if not html_content:
return None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_content.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
result = md_converter.convert(html_stream)
markdown_content = result.text_content
logger.info("Successfully converted HTML to Markdown")
return markdown_content
except Exception as e:
logger.error(f"Error converting HTML to Markdown: {e}")
return f"Error converting HTML content: {str(e)}"
def _convert_pdf_to_markdown(self, pdf_bytes: bytes) -> Optional[str]:
"""Convert PDF to Markdown using MarkItDown"""
if not pdf_bytes:
return None
try:
# Create BytesIO stream from PDF bytes
pdf_stream = io.BytesIO(pdf_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
result = md_converter.convert(pdf_stream)
markdown_content = result.text_content
logger.info("Successfully converted PDF to Markdown")
return markdown_content
except Exception as e:
logger.error(f"Error converting PDF to Markdown: {e}")
return f"Error converting PDF content: {str(e)}. The document may be corrupted or in an unsupported format."
async def close_client_session(self):
"""Close HTTP client session"""
await self.http_client.aclose()
logger.info("BedestenApiClient: HTTP client session closed.")
@@ -1,113 +0,0 @@
# bedesten_mcp_module/enums.py
from typing import Literal
# Unified compressed enum for both Yargıtay and Danıştay chambers
BirimAdiEnum = Literal[
"ALL", # All chambers
# Yargıtay (Court of Cassation) - Civil Chambers
"H1", "H2", "H3", "H4", "H5", "H6", "H7", "H8", "H9", "H10",
"H11", "H12", "H13", "H14", "H15", "H16", "H17", "H18", "H19", "H20",
"H21", "H22", "H23",
# Yargıtay - Criminal Chambers
"C1", "C2", "C3", "C4", "C5", "C6", "C7", "C8", "C9", "C10",
"C11", "C12", "C13", "C14", "C15", "C16", "C17", "C18", "C19", "C20",
"C21", "C22", "C23",
# Yargıtay - Councils and Assemblies
"HGK", # Hukuk Genel Kurulu
"CGK", # Ceza Genel Kurulu
"BGK", # Büyük Genel Kurulu
"HBK", # Hukuk Daireleri Başkanlar Kurulu
"CBK", # Ceza Daireleri Başkanlar Kurulu
# Danıştay (Council of State) - Chambers
"D1", "D2", "D3", "D4", "D5", "D6", "D7", "D8", "D9", "D10",
"D11", "D12", "D13", "D14", "D15", "D16", "D17",
# Danıştay - Councils and Boards
"DBGK", # Büyük Gen.Kur. (Grand General Assembly)
"IDDK", # İdare Dava Daireleri Kurulu
"VDDK", # Vergi Dava Daireleri Kurulu
"IBK", # İçtihatları Birleştirme Kurulu
"IIK", # İdari İşler Kurulu
"DBK", # Başkanlar Kurulu
# Military High Administrative Court
"AYIM", # Askeri Yüksek İdare Mahkemesi
"AYIMDK", # Askeri Yüksek İdare Mahkemesi Daireler Kurulu
"AYIMB", # Askeri Yüksek İdare Mahkemesi Başsavcılığı
"AYIM1", # Askeri Yüksek İdare Mahkemesi 1. Daire
"AYIM2", # Askeri Yüksek İdare Mahkemesi 2. Daire
"AYIM3" # Askeri Yüksek İdare Mahkemesi 3. Daire
]
# Mapping from abbreviated values to full Turkish API values
BIRIM_ADI_MAPPING = {
"ALL": None, # Will be handled specially in client
# Yargıtay Civil Chambers (1-23)
"H1": "1. Hukuk Dairesi", "H2": "2. Hukuk Dairesi", "H3": "3. Hukuk Dairesi",
"H4": "4. Hukuk Dairesi", "H5": "5. Hukuk Dairesi", "H6": "6. Hukuk Dairesi",
"H7": "7. Hukuk Dairesi", "H8": "8. Hukuk Dairesi", "H9": "9. Hukuk Dairesi",
"H10": "10. Hukuk Dairesi", "H11": "11. Hukuk Dairesi", "H12": "12. Hukuk Dairesi",
"H13": "13. Hukuk Dairesi", "H14": "14. Hukuk Dairesi", "H15": "15. Hukuk Dairesi",
"H16": "16. Hukuk Dairesi", "H17": "17. Hukuk Dairesi", "H18": "18. Hukuk Dairesi",
"H19": "19. Hukuk Dairesi", "H20": "20. Hukuk Dairesi", "H21": "21. Hukuk Dairesi",
"H22": "22. Hukuk Dairesi", "H23": "23. Hukuk Dairesi",
# Yargıtay Criminal Chambers (1-23)
"C1": "1. Ceza Dairesi", "C2": "2. Ceza Dairesi", "C3": "3. Ceza Dairesi",
"C4": "4. Ceza Dairesi", "C5": "5. Ceza Dairesi", "C6": "6. Ceza Dairesi",
"C7": "7. Ceza Dairesi", "C8": "8. Ceza Dairesi", "C9": "9. Ceza Dairesi",
"C10": "10. Ceza Dairesi", "C11": "11. Ceza Dairesi", "C12": "12. Ceza Dairesi",
"C13": "13. Ceza Dairesi", "C14": "14. Ceza Dairesi", "C15": "15. Ceza Dairesi",
"C16": "16. Ceza Dairesi", "C17": "17. Ceza Dairesi", "C18": "18. Ceza Dairesi",
"C19": "19. Ceza Dairesi", "C20": "20. Ceza Dairesi", "C21": "21. Ceza Dairesi",
"C22": "22. Ceza Dairesi", "C23": "23. Ceza Dairesi",
# Yargıtay Councils and Assemblies
"HGK": "Hukuk Genel Kurulu",
"CGK": "Ceza Genel Kurulu",
"BGK": "Büyük Genel Kurulu",
"HBK": "Hukuk Daireleri Başkanlar Kurulu",
"CBK": "Ceza Daireleri Başkanlar Kurulu",
# Danıştay Chambers (1-17)
"D1": "1. Daire", "D2": "2. Daire", "D3": "3. Daire", "D4": "4. Daire",
"D5": "5. Daire", "D6": "6. Daire", "D7": "7. Daire", "D8": "8. Daire",
"D9": "9. Daire", "D10": "10. Daire", "D11": "11. Daire", "D12": "12. Daire",
"D13": "13. Daire", "D14": "14. Daire", "D15": "15. Daire", "D16": "16. Daire",
"D17": "17. Daire",
# Danıştay Councils and Boards
"DBGK": "Büyük Gen.Kur.",
"IDDK": "İdare Dava Daireleri Kurulu",
"VDDK": "Vergi Dava Daireleri Kurulu",
"IBK": "İçtihatları Birleştirme Kurulu",
"IIK": "İdari İşler Kurulu",
"DBK": "Başkanlar Kurulu",
# Military High Administrative Court
"AYIM": "Askeri Yüksek İdare Mahkemesi",
"AYIMDK": "Askeri Yüksek İdare Mahkemesi Daireler Kurulu",
"AYIMB": "Askeri Yüksek İdare Mahkemesi Başsavcılığı",
"AYIM1": "Askeri Yüksek İdare Mahkemesi 1. Daire",
"AYIM2": "Askeri Yüksek İdare Mahkemesi 2. Daire",
"AYIM3": "Askeri Yüksek İdare Mahkemesi 3. Daire"
}
# Helper function to get full Turkish name from abbreviated value
def get_full_birim_adi(abbreviated_value: str) -> str:
"""Convert abbreviated birimAdi value to full Turkish name for API calls."""
if abbreviated_value == "ALL" or not abbreviated_value:
return "" # Empty string for ALL or None
return BIRIM_ADI_MAPPING.get(abbreviated_value, abbreviated_value)
# Helper function to validate abbreviated value
def is_valid_birim_adi(abbreviated_value: str) -> bool:
"""Check if abbreviated birimAdi value is valid."""
return abbreviated_value in BIRIM_ADI_MAPPING
@@ -1,91 +0,0 @@
# bedesten_mcp_module/models.py
from pydantic import BaseModel, Field
from typing import List, Optional, Dict, Any, Literal, Union
from datetime import datetime
# Import compressed BirimAdiEnum for chamber filtering
from .enums import BirimAdiEnum
# Court Type Options for Unified Search
BedestenCourtTypeEnum = Literal[
"YARGITAYKARARI", # Yargıtay (Court of Cassation)
"DANISTAYKARAR", # Danıştay (Council of State)
"YERELHUKUK", # Local Civil Courts
"ISTINAFHUKUK", # Civil Courts of Appeals
"KYB" # Extraordinary Appeals (Kanun Yararına Bozma)
]
# Search Request Models
class BedestenSearchData(BaseModel):
pageSize: int = Field(..., description="Results per page (1-10)")
pageNumber: int = Field(..., description="Page number (1-indexed)")
itemTypeList: List[str] = Field(..., description="Court type filter (YARGITAYKARARI/DANISTAYKARAR/YERELHUKUK/ISTINAFHUKUK/KYB)")
phrase: str = Field(..., description="Search phrase. Supports: 'word', \"exact phrase\", +required, -exclude, AND/OR/NOT operators. No wildcards or regex.")
birimAdi: BirimAdiEnum = Field("ALL", description="""
Chamber filter (optional). Abbreviated values with Turkish names:
Yargıtay: H1-H23 (1-23. Hukuk Dairesi), C1-C23 (1-23. Ceza Dairesi), HGK (Hukuk Genel Kurulu), CGK (Ceza Genel Kurulu), BGK (Büyük Genel Kurulu), HBK (Hukuk Daireleri Başkanlar Kurulu), CBK (Ceza Daireleri Başkanlar Kurulu)
Danıştay: D1-D17 (1-17. Daire), DBGK (Büyük Gen.Kur.), IDDK (İdare Dava Daireleri Kurulu), VDDK (Vergi Dava Daireleri Kurulu), IBK (İçtihatları Birleştirme Kurulu), IIK (İdari İşler Kurulu), DBK (Başkanlar Kurulu), AYIM (Askeri Yüksek İdare Mahkemesi), AYIM1-3 (Askeri Yüksek İdare Mahkemesi 1-3. Daire)
""")
kararTarihiStart: Optional[str] = Field(None, description="Start date (ISO 8601 format)")
kararTarihiEnd: Optional[str] = Field(None, description="End date (ISO 8601 format)")
sortFields: List[str] = Field(default=["KARAR_TARIHI"], description="Sort fields")
sortDirection: str = Field(default="desc", description="Sort direction (asc/desc)")
class BedestenSearchRequest(BaseModel):
data: BedestenSearchData
applicationName: str = "UyapMevzuat"
paging: bool = True
# Search Response Models
class BedestenItemType(BaseModel):
name: str
description: str
class BedestenDecisionEntry(BaseModel):
documentId: str
itemType: BedestenItemType
birimId: Optional[str] = None
birimAdi: Optional[str]
esasNoYil: Optional[int] = None
esasNoSira: Optional[int] = None
kararNoYil: Optional[int] = None
kararNoSira: Optional[int] = None
kararTuru: Optional[str] = None
kararTarihi: str
kararTarihiStr: str
kesinlesmeDurumu: Optional[str] = None
kararNo: Optional[str] = None
esasNo: Optional[str] = None
class BedestenSearchDataResponse(BaseModel):
emsalKararList: List[BedestenDecisionEntry]
total: int
start: int
class BedestenSearchResponse(BaseModel):
data: Optional[BedestenSearchDataResponse]
metadata: Dict[str, Any]
# Document Request/Response Models
class BedestenDocumentRequestData(BaseModel):
documentId: str
class BedestenDocumentRequest(BaseModel):
data: BedestenDocumentRequestData
applicationName: str = "UyapMevzuat"
class BedestenDocumentData(BaseModel):
content: str # Base64 encoded HTML or PDF
mimeType: str
version: int
class BedestenDocumentResponse(BaseModel):
data: BedestenDocumentData
metadata: Dict[str, Any]
class BedestenDocumentMarkdown(BaseModel):
documentId: str = Field(..., description="The document ID (Belge Kimliği) from Bedesten")
markdown_content: Optional[str] = Field(None, description="The decision content (Karar İçeriği) converted to Markdown")
source_url: str = Field(..., description="The source URL (Kaynak URL) of the document")
mime_type: Optional[str] = Field(None, description="Original content type (İçerik Türü) (text/html or application/pdf)")
@@ -1,21 +0,0 @@
#!/usr/bin/env python3
from fastmcp import Client
from mcp_server_main import app
import json
import asyncio
async def check_response_format():
client = Client(app)
async with client:
result = await client.call_tool('search_bedesten_unified', {
'phrase': 'mülkiyet',
'court_types': ['YARGITAYKARARI'],
'birimAdi': 'H1',
'pageSize': 3
})
if result and result.content:
data = json.loads(result.content[0].text)
print('Response keys:', list(data.keys()))
print('Sample response:', json.dumps(data, indent=2, ensure_ascii=False)[:500])
asyncio.run(check_response_format())
@@ -1,192 +0,0 @@
# danistay_mcp_module/client.py
import httpx
from bs4 import BeautifulSoup
from typing import Dict, Any, List, Optional
import logging
import html
import re
import io
from markitdown import MarkItDown
from .models import (
DanistayKeywordSearchRequest,
DanistayDetailedSearchRequest,
DanistayApiResponse,
DanistayDocumentMarkdown,
DanistayKeywordSearchRequestData,
DanistayDetailedSearchRequestData
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
class DanistayApiClient:
BASE_URL = "https://karararama.danistay.gov.tr"
KEYWORD_SEARCH_ENDPOINT = "/aramalist"
DETAILED_SEARCH_ENDPOINT = "/aramadetaylist"
DOCUMENT_ENDPOINT = "/getDokuman"
def __init__(self, request_timeout: float = 30.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Content-Type": "application/json; charset=UTF-8", # Arama endpoint'leri için
"Accept": "application/json, text/plain, */*", # Arama endpoint'leri için
"X-Requested-With": "XMLHttpRequest",
},
timeout=request_timeout,
verify=False
)
def _prepare_keywords_for_api(self, keywords: List[str]) -> List[str]:
return ['"' + k.strip('"') + '"' for k in keywords if k and k.strip()]
async def search_keyword_decisions(
self,
params: DanistayKeywordSearchRequest
) -> DanistayApiResponse:
data_for_payload = DanistayKeywordSearchRequestData(
andKelimeler=self._prepare_keywords_for_api(params.andKelimeler),
orKelimeler=self._prepare_keywords_for_api(params.orKelimeler),
notAndKelimeler=self._prepare_keywords_for_api(params.notAndKelimeler),
notOrKelimeler=self._prepare_keywords_for_api(params.notOrKelimeler),
pageSize=params.pageSize,
pageNumber=params.pageNumber
)
final_payload = {"data": data_for_payload.model_dump(exclude_none=True)}
logger.info(f"DanistayApiClient: Performing KEYWORD search via {self.KEYWORD_SEARCH_ENDPOINT} with payload: {final_payload}")
return await self._execute_api_search(self.KEYWORD_SEARCH_ENDPOINT, final_payload)
async def search_detailed_decisions(
self,
params: DanistayDetailedSearchRequest
) -> DanistayApiResponse:
data_for_payload = DanistayDetailedSearchRequestData(
daire=params.daire or "",
esasYil=params.esasYil or "",
esasIlkSiraNo=params.esasIlkSiraNo or "",
esasSonSiraNo=params.esasSonSiraNo or "",
kararYil=params.kararYil or "",
kararIlkSiraNo=params.kararIlkSiraNo or "",
kararSonSiraNo=params.kararSonSiraNo or "",
baslangicTarihi=params.baslangicTarihi or "",
bitisTarihi=params.bitisTarihi or "",
mevzuatNumarasi=params.mevzuatNumarasi or "",
mevzuatAdi=params.mevzuatAdi or "",
madde=params.madde or "",
siralama="1",
siralamaDirection="desc",
pageSize=params.pageSize,
pageNumber=params.pageNumber
)
# Create request dict and remove empty string fields to avoid API issues
payload_dict = data_for_payload.model_dump(exclude_defaults=False, exclude_none=False)
# Remove empty string fields that might cause API issues
cleaned_payload = {k: v for k, v in payload_dict.items() if v != ""}
final_payload = {"data": cleaned_payload}
logger.info(f"DanistayApiClient: Performing DETAILED search via {self.DETAILED_SEARCH_ENDPOINT} with payload: {final_payload}")
return await self._execute_api_search(self.DETAILED_SEARCH_ENDPOINT, final_payload)
async def _execute_api_search(self, endpoint: str, payload: Dict) -> DanistayApiResponse:
try:
response = await self.http_client.post(endpoint, json=payload)
response.raise_for_status()
response_json_data = response.json()
logger.debug(f"DanistayApiClient: Raw API response from {endpoint}: {response_json_data}")
api_response_parsed = DanistayApiResponse(**response_json_data)
if api_response_parsed.data and api_response_parsed.data.data:
for decision_item in api_response_parsed.data.data:
if decision_item.id:
decision_item.document_url = f"{self.BASE_URL}{self.DOCUMENT_ENDPOINT}?id={decision_item.id}"
return api_response_parsed
except httpx.RequestError as e:
logger.error(f"DanistayApiClient: HTTP request error during search to {endpoint}: {e}")
raise
except Exception as e:
logger.error(f"DanistayApiClient: Error processing or validating search response from {endpoint}: {e}")
raise
def _convert_html_to_markdown_danistay(self, direct_html_content: str) -> Optional[str]:
"""
Converts direct HTML content (assumed from Danıştay /getDokuman) to Markdown.
"""
if not direct_html_content:
return None
# Basic HTML unescaping and fixing common escaped characters
# This step might be less critical if MarkItDown handles them, but good for pre-cleaning.
processed_html = html.unescape(direct_html_content)
processed_html = processed_html.replace('\\"', '"') # If any such JS-escaped strings exist
# Danistay HTML doesn't seem to have \\r\\n etc. from the example, but keeping for robustness
processed_html = processed_html.replace('\\r\\n', '\n').replace('\\n', '\n').replace('\\t', '\t')
# For simplicity and to leverage MarkItDown's capability to handle full docs,
# we pass the pre-processed full HTML.
html_input_for_markdown = processed_html
markdown_text = None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_input_for_markdown.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_text = conversion_result.text_content
logger.info("DanistayApiClient: HTML to Markdown conversion successful.")
except Exception as e:
logger.error(f"DanistayApiClient: Error during MarkItDown HTML to Markdown conversion: {e}")
return markdown_text
async def get_decision_document_as_markdown(self, id: str) -> DanistayDocumentMarkdown:
"""
Retrieves a specific Danıştay decision by ID and returns its content as Markdown.
The /getDokuman endpoint for Danıştay requires arananKelime parameter.
"""
# Add required arananKelime parameter - using empty string as minimum requirement
document_api_url = f"{self.DOCUMENT_ENDPOINT}?id={id}&arananKelime="
source_url = f"{self.BASE_URL}{document_api_url}"
logger.info(f"DanistayApiClient: Fetching Danistay document for Markdown (ID: {id}) from {source_url}")
try:
# For direct HTML response, we might want different headers if the API is sensitive,
# but httpx usually handles basic GET requests well.
response = await self.http_client.get(document_api_url)
response.raise_for_status()
# Danıştay /getDokuman directly returns HTML text
html_content_from_api = response.text
if not isinstance(html_content_from_api, str) or not html_content_from_api.strip():
logger.warning(f"DanistayApiClient: Received empty or non-string HTML content for ID {id}.")
# Return with None markdown_content if HTML is effectively empty
return DanistayDocumentMarkdown(
id=id,
markdown_content=None,
source_url=source_url
)
markdown_content = self._convert_html_to_markdown_danistay(html_content_from_api)
return DanistayDocumentMarkdown(
id=id,
markdown_content=markdown_content,
source_url=source_url
)
except httpx.RequestError as e:
logger.error(f"DanistayApiClient: HTTP error fetching Danistay document (ID: {id}): {e}")
raise
# Removed ValueError for JSON as Danistay /getDokuman returns direct HTML
except Exception as e: # Catches other errors like MarkItDown issues if they propagate
logger.error(f"DanistayApiClient: General error processing Danistay document (ID: {id}): {e}")
raise
async def close_client_session(self):
"""Closes the HTTPX client session."""
if self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("DanistayApiClient: HTTP client session closed.")
@@ -1,112 +0,0 @@
# danistay_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl, ConfigDict
from typing import List, Optional, Dict, Any
class DanistayBaseSearchRequest(BaseModel):
"""Base model for common search parameters for Danistay."""
pageSize: int = Field(default=10, ge=1, le=10)
pageNumber: int = Field(default=1, ge=1)
# siralama and siralamaDirection are part of detailed search, not necessarily keyword search
# as per user's provided payloads.
class DanistayKeywordSearchRequestData(BaseModel):
"""Internal data model for the keyword search payload's 'data' field."""
andKelimeler: List[str] = Field(default_factory=list)
orKelimeler: List[str] = Field(default_factory=list)
notAndKelimeler: List[str] = Field(default_factory=list)
notOrKelimeler: List[str] = Field(default_factory=list)
pageSize: int
pageNumber: int
class DanistayKeywordSearchRequest(BaseModel): # This is the model the MCP tool will accept
"""Model for keyword-based search request for Danistay."""
andKelimeler: List[str] = Field(default_factory=list, description="AND keywords")
orKelimeler: List[str] = Field(default_factory=list, description="OR keywords")
notAndKelimeler: List[str] = Field(default_factory=list, description="NOT AND keywords")
notOrKelimeler: List[str] = Field(default_factory=list, description="NOT OR keywords")
pageSize: int = Field(default=10, ge=1, le=10)
pageNumber: int = Field(default=1, ge=1)
class DanistayDetailedSearchRequestData(BaseModel): # Internal data model for detailed search payload
"""Internal data model for the detailed search payload's 'data' field."""
daire: Optional[str] = "" # API expects empty string for None
esasYil: Optional[str] = ""
esasIlkSiraNo: Optional[str] = ""
esasSonSiraNo: Optional[str] = ""
kararYil: Optional[str] = ""
kararIlkSiraNo: Optional[str] = ""
kararSonSiraNo: Optional[str] = ""
baslangicTarihi: Optional[str] = ""
bitisTarihi: Optional[str] = ""
mevzuatNumarasi: Optional[str] = ""
mevzuatAdi: Optional[str] = ""
madde: Optional[str] = ""
siralama: str # Seems mandatory in detailed search payload
siralamaDirection: str # Seems mandatory
pageSize: int
pageNumber: int
# Note: 'arananKelime' is not in the detailed search payload example provided by user.
# If it can be included, it should be added here.
class DanistayDetailedSearchRequest(DanistayBaseSearchRequest): # MCP tool will accept this
"""Model for detailed search request for Danistay."""
daire: str = Field("", description="Chamber")
esasYil: str = Field("", description="Case year")
esasIlkSiraNo: str = Field("", description="Start case no")
esasSonSiraNo: str = Field("", description="End case no")
kararYil: str = Field("", description="Decision year")
kararIlkSiraNo: str = Field("", description="Start decision no")
kararSonSiraNo: str = Field("", description="End decision no")
baslangicTarihi: str = Field("", description="Start date")
bitisTarihi: str = Field("", description="End date")
mevzuatNumarasi: str = Field("", description="Law number")
mevzuatAdi: str = Field("", description="Law name")
madde: str = Field("", description="Article")
# Add a general keyword field if detailed search also supports it
# arananKelime: Optional[str] = Field(None, description="General keyword for detailed search.")
class DanistayApiDecisionEntry(BaseModel):
"""Model for an individual decision entry from the Danistay API search response.
Based on user-provided response samples for both keyword and detailed search.
"""
id: str
# The API response for keyword search uses "daireKurul", detailed search example uses "daire".
# We use an alias to handle both and map to a consistent field name "chamber".
chamber: str = Field("", alias="daire", description="Chamber")
esasNo: str = Field("", description="Case number")
kararNo: str = Field("", description="Decision number")
kararTarihi: str = Field("", description="Decision date")
arananKelime: str = Field("", description="Keyword")
# index: Optional[int] = None # Present in response, can be added if needed by MCP tool
# siraNo: Optional[int] = None # Present in detailed response, can be added
document_url: Optional[HttpUrl] = Field(None, description="Document URL")
model_config = ConfigDict(populate_by_name=True, extra='ignore') # Important for alias to work and ignore extra fields
class DanistayApiResponseInnerData(BaseModel):
"""Model for the inner 'data' object in the Danistay API search response."""
data: List[DanistayApiDecisionEntry]
recordsTotal: int
recordsFiltered: int
draw: int = Field(0, description="Draw counter")
class DanistayApiResponse(BaseModel):
"""Model for the complete search response from the Danistay API."""
data: Optional[DanistayApiResponseInnerData] = Field(None, description="Response data, can be null when no results found")
metadata: Optional[Dict[str, Any]] = Field(None, description="Optional metadata (Meta Veri) from API.")
class DanistayDocumentMarkdown(BaseModel):
"""Model for a Danistay decision document, containing only Markdown content."""
id: str
markdown_content: str = Field("", description="The decision content (Karar İçeriği) converted to Markdown.")
source_url: HttpUrl
class CompactDanistaySearchResult(BaseModel):
"""A compact search result model for the MCP tool to return."""
decisions: List[DanistayApiDecisionEntry]
total_records: int
requested_page: int
page_size: int
@@ -1,66 +0,0 @@
version: '3.8'
services:
yargi-mcp:
build: .
image: yargi-mcp:latest
container_name: yargi-mcp-server
ports:
- "${PORT:-8000}:8000"
environment:
- HOST=0.0.0.0
- PORT=8000
- LOG_LEVEL=${LOG_LEVEL:-info}
- ALLOWED_ORIGINS=${ALLOWED_ORIGINS:-*}
- API_TOKEN=${API_TOKEN:-}
- PYTHONUNBUFFERED=1
volumes:
# Mount logs directory
- ./logs:/app/logs
# Mount .env file if it exists
- ./.env:/app/.env:ro
restart: unless-stopped
healthcheck:
test: ["CMD", "python", "-c", "import httpx; httpx.get('http://localhost:8000/health').raise_for_status()"]
interval: 30s
timeout: 10s
retries: 3
start_period: 10s
networks:
- yargi-network
# Optional: Nginx reverse proxy
nginx:
image: nginx:alpine
container_name: yargi-nginx
ports:
- "80:80"
- "443:443"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf:ro
- ./ssl:/etc/nginx/ssl:ro
depends_on:
- yargi-mcp
networks:
- yargi-network
profiles:
- production
# Optional: Redis for caching (future enhancement)
redis:
image: redis:alpine
container_name: yargi-redis
command: redis-server --appendonly yes
volumes:
- redis-data:/data
networks:
- yargi-network
profiles:
- with-cache
networks:
yargi-network:
driver: bridge
volumes:
redis-data:
@@ -1,428 +0,0 @@
# Yargı MCP Server Dağıtım Rehberi
Bu rehber, Yargı MCP Server'ın ASGI web servisi olarak çeşitli dağıtım seçeneklerini kapsar.
## İçindekiler
- [Hızlı Başlangıç](#hızlı-başlangıç)
- [Yerel Geliştirme](#yerel-geliştirme)
- [Production Dağıtımı](#production-dağıtımı)
- [Cloud Dağıtımı](#cloud-dağıtımı)
- [Docker Dağıtımı](#docker-dağıtımı)
- [Güvenlik Hususları](#güvenlik-hususları)
- [İzleme](#izleme)
## Hızlı Başlangıç
### 1. Bağımlılıkları Yükleyin
```bash
# ASGI sunucusu için uvicorn yükleyin
pip install uvicorn
# Veya tüm bağımlılıklarla birlikte yükleyin
pip install -e .
pip install uvicorn
```
### 2. Sunucuyu Çalıştırın
```bash
# Temel başlatma
python run_asgi.py
# Veya doğrudan uvicorn ile
uvicorn asgi_app:app --host 0.0.0.0 --port 8000
```
Sunucu şu adreslerde kullanılabilir olacak:
- MCP Endpoint: `http://localhost:8000/mcp/`
- Sağlık Kontrolü: `http://localhost:8000/health`
- API Durumu: `http://localhost:8000/status`
## Yerel Geliştirme
### Otomatik Yeniden Yükleme ile Geliştirme Sunucusu
```bash
python run_asgi.py --reload --log-level debug
```
### FastAPI Entegrasyonunu Kullanma
Ek REST API endpoint'leri için:
```bash
uvicorn fastapi_app:app --reload
```
Bu şunları sağlar:
- `/docs` adresinde interaktif API dokümantasyonu
- `/api/tools` adresinde araç listesi
- `/api/databases` adresinde veritabanı bilgileri
### Ortam Değişkenleri
`.env.example` dosyasını temel alarak bir `.env` dosyası oluşturun:
```bash
cp .env.example .env
```
Temel değişkenler:
- `HOST`: Sunucu host adresi (varsayılan: 127.0.0.1)
- `PORT`: Sunucu portu (varsayılan: 8000)
- `ALLOWED_ORIGINS`: CORS kökenleri (virgülle ayrılmış)
- `LOG_LEVEL`: Log seviyesi (debug, info, warning, error)
## Production Dağıtımı
### 1. Uvicorn ile Çoklu Worker Kullanımı
```bash
python run_asgi.py --host 0.0.0.0 --port 8000 --workers 4
```
### 2. Gunicorn Kullanımı
```bash
pip install gunicorn
gunicorn asgi_app:app -w 4 -k uvicorn.workers.UvicornWorker --bind 0.0.0.0:8000
```
### 3. Nginx Reverse Proxy ile
1. Nginx'i yükleyin
2. Sağlanan `nginx.conf` dosyasını kullanın:
```bash
sudo cp nginx.conf /etc/nginx/sites-available/yargi-mcp
sudo ln -s /etc/nginx/sites-available/yargi-mcp /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl reload nginx
```
### 4. Systemd Servisi
`/etc/systemd/system/yargi-mcp.service` dosyasını oluşturun:
```ini
[Unit]
Description=Yargı MCP Server
After=network.target
[Service]
Type=exec
User=www-data
WorkingDirectory=/opt/yargi-mcp
Environment="PATH=/opt/yargi-mcp/venv/bin"
ExecStart=/opt/yargi-mcp/venv/bin/uvicorn asgi_app:app --host 0.0.0.0 --port 8000 --workers 4
Restart=on-failure
RestartSec=5
[Install]
WantedBy=multi-user.target
```
Etkinleştirin ve başlatın:
```bash
sudo systemctl enable yargi-mcp
sudo systemctl start yargi-mcp
```
## Cloud Dağıtımı
### Heroku
1. `Procfile` oluşturun:
```
web: uvicorn asgi_app:app --host 0.0.0.0 --port $PORT
```
2. Dağıtın:
```bash
heroku create uygulama-isminiz
git push heroku main
```
### Railway
1. `railway.json` ekleyin:
```json
{
"build": {
"builder": "NIXPACKS"
},
"deploy": {
"startCommand": "uvicorn asgi_app:app --host 0.0.0.0 --port $PORT"
}
}
```
2. Railway CLI veya GitHub entegrasyonu ile dağıtın
### Google Cloud Run
1. Container oluşturun:
```bash
docker build -t yargi-mcp .
docker tag yargi-mcp gcr.io/PROJE_ADINIZ/yargi-mcp
docker push gcr.io/PROJE_ADINIZ/yargi-mcp
```
2. Dağıtın:
```bash
gcloud run deploy yargi-mcp \
--image gcr.io/PROJE_ADINIZ/yargi-mcp \
--platform managed \
--region us-central1 \
--allow-unauthenticated
```
### AWS Lambda (Mangum kullanarak)
1. Mangum'u yükleyin:
```bash
pip install mangum
```
2. `lambda_handler.py` oluşturun:
```python
from mangum import Mangum
from asgi_app import app
handler = Mangum(app, lifespan="off")
```
3. AWS SAM veya Serverless Framework kullanarak dağıtın
## Docker Dağıtımı
### Tek Container
```bash
# Oluşturun
docker build -t yargi-mcp .
# Çalıştırın
docker run -p 8000:8000 --env-file .env yargi-mcp
```
### Docker Compose
```bash
# Geliştirme
docker-compose up
# Nginx ile Production
docker-compose --profile production up
# Redis önbellekleme ile
docker-compose --profile with-cache up
```
### Kubernetes
Deployment YAML oluşturun:
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: yargi-mcp
spec:
replicas: 3
selector:
matchLabels:
app: yargi-mcp
template:
metadata:
labels:
app: yargi-mcp
spec:
containers:
- name: yargi-mcp
image: yargi-mcp:latest
ports:
- containerPort: 8000
env:
- name: HOST
value: "0.0.0.0"
- name: PORT
value: "8000"
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 10
periodSeconds: 30
---
apiVersion: v1
kind: Service
metadata:
name: yargi-mcp-service
spec:
selector:
app: yargi-mcp
ports:
- port: 80
targetPort: 8000
type: LoadBalancer
```
## Güvenlik Hususları
### 1. Kimlik Doğrulama
`API_TOKEN` ortam değişkenini ayarlayarak token kimlik doğrulamasını etkinleştirin:
```bash
export API_TOKEN=gizli-token-degeri
```
Ardından isteklere ekleyin:
```bash
curl -H "Authorization: Bearer gizli-token-degeri" http://localhost:8000/api/tools
```
### 2. HTTPS/SSL
Production için her zaman HTTPS kullanın:
1. SSL sertifikası edinin (Let's Encrypt vb.)
2. Nginx veya cloud sağlayıcıda yapılandırın
3. `ALLOWED_ORIGINS` değerini https:// kullanacak şekilde güncelleyin
### 3. Rate Limiting (Hız Sınırlama)
Sağlanan Nginx yapılandırması rate limiting içerir:
- API endpoint'leri: 10 istek/saniye
- MCP endpoint: 100 istek/saniye
### 4. CORS Yapılandırması
Production için belirli kaynaklara izin verin:
```bash
ALLOWED_ORIGINS=https://app.sizindomain.com,https://www.sizindomain.com
```
## İzleme
### Sağlık Kontrolleri
`/health` endpoint'ini izleyin:
```bash
curl http://localhost:8000/health
```
Yanıt:
```json
{
"status": "healthy",
"timestamp": "2024-12-26T10:00:00",
"uptime_seconds": 3600,
"tools_operational": true
}
```
### Loglama
Ortam değişkeni ile log seviyesini yapılandırın:
```bash
LOG_LEVEL=info # veya debug, warning, error
```
Loglar şuraya yazılır:
- Konsol (stdout)
- `logs/mcp_server.log` dosyası
### Metrikler (Opsiyonel)
OpenTelemetry desteği için:
```bash
pip install opentelemetry-instrumentation-fastapi
```
Ortam değişkenlerini ayarlayın:
```bash
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
OTEL_SERVICE_NAME=yargi-mcp-server
```
## Sorun Giderme
### Port Zaten Kullanımda
```bash
# 8000 portunu kullanan işlemi bulun
lsof -i :8000
# İşlemi sonlandırın
kill -9 <PID>
```
### İzin Hataları
Dosya izinlerinin doğru olduğundan emin olun:
```bash
chmod +x run_asgi.py
chown -R www-data:www-data /opt/yargi-mcp
```
### Bellek Sorunları
Büyük belge işleme için worker belleğini artırın:
```bash
# systemd servisinde
Environment="PYTHONMALLOC=malloc"
LimitNOFILE=65536
```
### Zaman Aşımı Sorunları
Zaman aşımlarını ayarlayın:
1. Uvicorn: `--timeout-keep-alive 75`
2. Nginx: `proxy_read_timeout 300s;`
3. Cloud sağlayıcılar: Platform özel zaman aşımı ayarlarını kontrol edin
## Performans Ayarlama
### 1. Worker İşlemleri
- Geliştirme: 1 worker
- Production: CPU çekirdeği başına 2-4 worker
### 2. Bağlantı Havuzlama
Sunucu varsayılan olarak httpx ile bağlantı havuzlama kullanır.
### 3. Önbellekleme (Gelecek Geliştirme)
Redis önbellekleme docker-compose ile etkinleştirilebilir:
```bash
docker-compose --profile with-cache up
```
### 4. Veritabanı Zaman Aşımları
`.env` dosyasında veritabanı başına zaman aşımlarını ayarlayın:
```bash
YARGITAY_TIMEOUT=60
DANISTAY_TIMEOUT=60
ANAYASA_TIMEOUT=90
```
## Destek
Sorunlar ve sorular için:
- GitHub Issues: https://github.com/saidsurucu/yargi-mcp/issues
- Dokümantasyon: README.md dosyasına bakın
@@ -1,177 +0,0 @@
# emsal_mcp_module/client.py
import httpx
# from bs4 import BeautifulSoup # Uncomment if needed for advanced HTML pre-processing
from typing import Dict, Any, List, Optional
import logging
import html
import re
import io
from markitdown import MarkItDown
from .models import (
EmsalSearchRequest,
EmsalDetailedSearchRequestData,
EmsalApiResponse,
EmsalDocumentMarkdown
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
class EmsalApiClient:
"""API Client for Emsal (UYAP Precedent Decision) search system."""
BASE_URL = "https://emsal.uyap.gov.tr"
DETAILED_SEARCH_ENDPOINT = "/aramadetaylist"
DOCUMENT_ENDPOINT = "/getDokuman"
def __init__(self, request_timeout: float = 30.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Content-Type": "application/json; charset=UTF-8",
"Accept": "application/json, text/plain, */*",
"X-Requested-With": "XMLHttpRequest",
},
timeout=request_timeout,
verify=False # As per user's original FastAPI code
)
async def search_detailed_decisions(
self,
params: EmsalSearchRequest
) -> EmsalApiResponse:
"""Performs a detailed search for Emsal decisions."""
data_for_api_payload = EmsalDetailedSearchRequestData(
arananKelime=params.keyword or "",
Bam_Hukuk_Mahkemeleri=params.selected_bam_civil_court, # Uses alias "Bam Hukuk Mahkemeleri"
Hukuk_Mahkemeleri=params.selected_civil_court, # Uses alias "Hukuk Mahkemeleri"
birimHukukMah="+".join(params.selected_regional_civil_chambers) if params.selected_regional_civil_chambers else "",
esasYil=params.case_year_esas or "",
esasIlkSiraNo=params.case_start_seq_esas or "",
esasSonSiraNo=params.case_end_seq_esas or "",
kararYil=params.decision_year_karar or "",
kararIlkSiraNo=params.decision_start_seq_karar or "",
kararSonSiraNo=params.decision_end_seq_karar or "",
baslangicTarihi=params.start_date or "",
bitisTarihi=params.end_date or "",
siralama=params.sort_criteria,
siralamaDirection=params.sort_direction,
pageSize=params.page_size,
pageNumber=params.page_number
)
# Create request dict and remove empty string fields to avoid API issues
payload_dict = data_for_api_payload.model_dump(by_alias=True, exclude_none=True)
# Remove empty string fields that might cause API issues
cleaned_payload = {k: v for k, v in payload_dict.items() if v != ""}
final_payload = {"data": cleaned_payload}
logger.info(f"EmsalApiClient: Performing DETAILED search with payload: {final_payload}")
return await self._execute_api_search(self.DETAILED_SEARCH_ENDPOINT, final_payload)
async def _execute_api_search(self, endpoint: str, payload: Dict) -> EmsalApiResponse:
"""Helper method to execute search POST request and process response for Emsal."""
try:
response = await self.http_client.post(endpoint, json=payload)
response.raise_for_status()
response_json_data = response.json()
logger.debug(f"EmsalApiClient: Raw API response from {endpoint}: {response_json_data}")
api_response_parsed = EmsalApiResponse(**response_json_data)
if api_response_parsed.data and api_response_parsed.data.data:
for decision_item in api_response_parsed.data.data:
if decision_item.id:
decision_item.document_url = f"{self.BASE_URL}{self.DOCUMENT_ENDPOINT}?id={decision_item.id}"
return api_response_parsed
except httpx.RequestError as e:
logger.error(f"EmsalApiClient: HTTP request error during Emsal search to {endpoint}: {e}")
raise
except Exception as e:
logger.error(f"EmsalApiClient: Error processing or validating Emsal search response from {endpoint}: {e}")
raise
def _clean_html_and_convert_to_markdown_emsal(self, html_content_from_api_data_field: str) -> Optional[str]:
"""
Cleans HTML (from Emsal API 'data' field containing HTML string)
and converts it to Markdown using MarkItDown.
This assumes Emsal /getDokuman response is JSON with HTML in "data" field,
similar to Yargitay and the last Emsal /getDokuman example.
"""
if not html_content_from_api_data_field:
return None
# Basic HTML unescaping and fixing common escaped characters
# Based on user's original fix_html_content in app/routers/emsal.py
content = html.unescape(html_content_from_api_data_field)
content = content.replace('\\"', '"')
content = content.replace('\\r\\n', '\n')
content = content.replace('\\n', '\n')
content = content.replace('\\t', '\t')
# The HTML string from "data" field starts with "<html><head>..."
html_input_for_markdown = content
markdown_text = None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_input_for_markdown.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_text = conversion_result.text_content
logger.info("EmsalApiClient: HTML to Markdown conversion successful.")
except Exception as e:
logger.error(f"EmsalApiClient: Error during MarkItDown HTML to Markdown conversion for Emsal: {e}")
return markdown_text
async def get_decision_document_as_markdown(self, id: str) -> EmsalDocumentMarkdown:
"""
Retrieves a specific Emsal decision by ID and returns its content as Markdown.
Assumes Emsal /getDokuman endpoint returns JSON with HTML content in the 'data' field.
"""
document_api_url = f"{self.DOCUMENT_ENDPOINT}?id={id}"
source_url = f"{self.BASE_URL}{document_api_url}"
logger.info(f"EmsalApiClient: Fetching Emsal document for Markdown (ID: {id}) from {source_url}")
try:
response = await self.http_client.get(document_api_url)
response.raise_for_status()
# Emsal /getDokuman returns JSON with HTML in 'data' field (confirmed by user example)
response_json = response.json()
html_content_from_api = response_json.get("data")
if not isinstance(html_content_from_api, str) or not html_content_from_api.strip():
logger.warning(f"EmsalApiClient: Received empty or non-string HTML in 'data' field for Emsal ID {id}.")
return EmsalDocumentMarkdown(id=id, markdown_content=None, source_url=source_url)
markdown_content = self._clean_html_and_convert_to_markdown_emsal(html_content_from_api)
return EmsalDocumentMarkdown(
id=id,
markdown_content=markdown_content,
source_url=source_url
)
except httpx.RequestError as e:
logger.error(f"EmsalApiClient: HTTP error fetching Emsal document (ID: {id}): {e}")
raise
except ValueError as e:
logger.error(f"EmsalApiClient: ValueError processing Emsal document response (ID: {id}): {e}")
raise
except Exception as e:
logger.error(f"EmsalApiClient: General error processing Emsal document (ID: {id}): {e}")
raise
async def close_client_session(self):
"""Closes the HTTPX client session."""
if self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("EmsalApiClient: HTTP client session closed.")
@@ -1,101 +0,0 @@
# emsal_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl, ConfigDict
from typing import List, Optional, Dict, Any
class EmsalDetailedSearchRequestData(BaseModel):
"""
Internal model for the 'data' object in the Emsal detailed search payload.
Field names use aliases to match the exact keys in the API payload
(e.g., "Bam Hukuk Mahkemeleri" with spaces).
The API expects empty strings for None/omitted optional fields.
"""
arananKelime: Optional[str] = ""
Bam_Hukuk_Mahkemeleri: str = Field("", alias="Bam Hukuk Mahkemeleri")
Hukuk_Mahkemeleri: str = Field("", alias="Hukuk Mahkemeleri")
# Add other specific court type fields from the form if they are separate keys in payload
# E.g., "Ceza Mahkemeleri", "İdari Mahkemeler" etc.
birimHukukMah: Optional[str] = Field("", description="Regional chambers (+ separated)")
esasYil: Optional[str] = ""
esasIlkSiraNo: Optional[str] = ""
esasSonSiraNo: Optional[str] = ""
kararYil: Optional[str] = ""
kararIlkSiraNo: Optional[str] = ""
kararSonSiraNo: Optional[str] = ""
baslangicTarihi: Optional[str] = ""
bitisTarihi: Optional[str] = ""
siralama: str # Mandatory in payload example
siralamaDirection: str # Mandatory in payload example
pageSize: int
pageNumber: int
model_config = ConfigDict(populate_by_name=True) # Enables use of alias in serialization (when dumping to dict for payload)
class EmsalSearchRequest(BaseModel): # This is the model the MCP tool will accept
"""Model for Emsal detailed search request, with user-friendly field names."""
keyword: str = Field("", description="Keyword")
selected_bam_civil_court: str = Field("", description="BAM Civil Court")
selected_civil_court: str = Field("", description="Civil Court")
selected_regional_civil_chambers: List[str] = Field(default_factory=list, description="Regional chambers")
case_year_esas: str = Field("", description="Case year")
case_start_seq_esas: str = Field("", description="Start case no")
case_end_seq_esas: str = Field("", description="End case no")
decision_year_karar: str = Field("", description="Decision year")
decision_start_seq_karar: str = Field("", description="Start decision no")
decision_end_seq_karar: str = Field("", description="End decision no")
start_date: str = Field("", description="Start date (DD.MM.YYYY)")
end_date: str = Field("", description="End date (DD.MM.YYYY)")
sort_criteria: str = Field("1", description="Sort by")
sort_direction: str = Field("desc", description="Direction")
page_number: int = Field(default=1, ge=1)
page_size: int = Field(default=10, ge=1, le=10)
class EmsalApiDecisionEntry(BaseModel):
"""Model for an individual decision entry from the Emsal API search response."""
id: str
daire: str = Field("", description="Chamber")
esasNo: str = Field("", description="Case number")
kararNo: str = Field("", description="Decision number")
kararTarihi: str = Field("", description="Decision date")
arananKelime: str = Field("", description="Keyword")
durum: str = Field("", description="Status")
# index: Optional[int] = None # Present in Emsal response, can be added if tool needs it
document_url: Optional[HttpUrl] = Field(None, description="Document URL")
model_config = ConfigDict(extra='ignore')
class EmsalApiResponseInnerData(BaseModel):
"""Model for the inner 'data' object in the Emsal API search response."""
data: List[EmsalApiDecisionEntry]
recordsTotal: int
recordsFiltered: int
draw: int = Field(0, description="Draw counter (Çizim Sayıcısı) from API, usually for DataTables.")
class EmsalApiResponse(BaseModel):
"""Model for the complete search response from the Emsal API."""
data: EmsalApiResponseInnerData
metadata: Optional[Dict[str, Any]] = Field(None, description="Optional metadata (Meta Veri) from API, if any.")
class EmsalDocumentMarkdown(BaseModel):
"""Model for an Emsal decision document, containing only Markdown content."""
id: str
markdown_content: str = Field("", description="The decision content (Karar İçeriği) converted to Markdown.")
source_url: HttpUrl
class CompactEmsalSearchResult(BaseModel):
"""A compact search result model for the MCP tool to return."""
decisions: List[EmsalApiDecisionEntry]
total_records: int
requested_page: int
page_size: int
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,74 +0,0 @@
# kik_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl, computed_field, ConfigDict
from typing import List, Optional
from enum import Enum
import base64 # Base64 encoding/decoding için
class KikKararTipi(str, Enum):
"""Enum for KIK (Public Procurement Authority) Decision Types."""
UYUSMAZLIK = "rbUyusmazlik"
DUZENLEYICI = "rbDuzenleyici"
MAHKEME = "rbMahkeme"
class KikSearchRequest(BaseModel):
"""Model for KIK Decision search criteria."""
karar_tipi: KikKararTipi = Field(KikKararTipi.UYUSMAZLIK, description="Type")
karar_no: str = Field("", description="No")
karar_tarihi_baslangic: str = Field("", description="Start", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
karar_tarihi_bitis: str = Field("", description="End", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
resmi_gazete_sayisi: str = Field("", description="Gazette")
resmi_gazete_tarihi: str = Field("", description="Date", pattern=r"^\d{2}\.\d{2}\.\d{4}$|^$")
basvuru_konusu_ihale: str = Field("", description="Subject")
basvuru_sahibi: str = Field("", description="Applicant")
ihaleyi_yapan_idare: str = Field("", description="Entity")
yil: str = Field("", description="Year")
karar_metni: str = Field("", description="Text")
page: int = Field(1, ge=1, description="Page")
class KikDecisionEntry(BaseModel):
"""Represents a single decision entry from KIK search results."""
preview_event_target: str = Field(..., description="Event target")
karar_no_str: str = Field(..., alias="kararNo", description="Decision number")
karar_tipi: KikKararTipi = Field(..., description="Decision type")
karar_tarihi_str: str = Field(..., alias="kararTarihi", description="Date")
idare_str: str = Field("", alias="idare", description="Entity")
basvuru_sahibi_str: str = Field("", alias="basvuruSahibi", description="Applicant")
ihale_konusu_str: str = Field("", alias="ihaleKonusu", description="Subject")
@computed_field
@property
def karar_id(self) -> str:
"""
A Base64 encoded unique ID for the decision, combining decision type and number.
Format before encoding: "{karar_tipi.value}|{karar_no_str}"
"""
combined_key = f"{self.karar_tipi.value}|{self.karar_no_str}"
return base64.b64encode(combined_key.encode('utf-8')).decode('utf-8')
model_config = ConfigDict(populate_by_name=True)
class KikSearchResult(BaseModel):
"""Model for KIK search results."""
decisions: List[KikDecisionEntry]
total_records: int = 0
current_page: int = 1
class KikDocumentMarkdown(BaseModel):
"""
KIK decision document, with Markdown content potentially paginated.
"""
retrieved_with_karar_id: Optional[str] = Field(None, description="Request ID")
retrieved_karar_no: Optional[str] = Field(None, description="Decision number")
retrieved_karar_tipi: Optional[KikKararTipi] = Field(None, description="Decision type")
karar_id_param_from_url: Optional[str] = Field(None, alias="kararIdParam", description="Internal ID")
markdown_chunk: Optional[str] = Field(None, description="Content")
source_url: Optional[str] = Field(None, description="Source URL")
error_message: Optional[str] = Field(None, description="Error")
current_page: int = Field(1, description="Page")
total_pages: int = Field(1, description="Total pages")
is_paginated: bool = Field(False, description="Paginated")
full_content_char_count: Optional[int] = Field(None, description="Char count")
model_config = ConfigDict(populate_by_name=True)
@@ -1 +0,0 @@
# kvkk_mcp_module/__init__.py
@@ -1,372 +0,0 @@
# kvkk_mcp_module/client.py
import httpx
from bs4 import BeautifulSoup
from typing import List, Optional, Dict, Any
import logging
import os
import re
import io
import math
from urllib.parse import urljoin, urlparse, parse_qs
from markitdown import MarkItDown
from pydantic import HttpUrl
from .models import (
KvkkSearchRequest,
KvkkDecisionSummary,
KvkkSearchResult,
KvkkDocumentMarkdown
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
class KvkkApiClient:
"""
API client for searching and retrieving KVKK (Personal Data Protection Authority) decisions
using Brave Search API for discovery and direct HTTP requests for content retrieval.
"""
BRAVE_API_URL = "https://api.search.brave.com/res/v1/web/search"
KVKK_BASE_URL = "https://www.kvkk.gov.tr"
DOCUMENT_MARKDOWN_CHUNK_SIZE = 5000 # Character limit per page
def __init__(self, request_timeout: float = 60.0):
"""Initialize the KVKK API client."""
self.brave_api_token = os.getenv("BRAVE_API_TOKEN")
if not self.brave_api_token:
# Fallback to provided free token
self.brave_api_token = "BSAuaRKB-dvSDSQxIN0ft1p2k6N82Kq"
logger.info("Using fallback Brave API token (limited free token)")
else:
logger.info("Using Brave API token from environment variable")
self.http_client = httpx.AsyncClient(
headers={
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
},
timeout=request_timeout,
verify=True,
follow_redirects=True
)
def _construct_search_query(self, keywords: str) -> str:
"""Construct the search query for Brave API."""
base_query = 'site:kvkk.gov.tr "karar özeti"'
if keywords.strip():
return f"{base_query} {keywords.strip()}"
return base_query
def _extract_decision_id_from_url(self, url: str) -> Optional[str]:
"""Extract decision ID from KVKK decision URL."""
try:
# Example URL: https://www.kvkk.gov.tr/Icerik/7288/2021-1303
parsed_url = urlparse(url)
path_parts = parsed_url.path.strip('/').split('/')
if len(path_parts) >= 3 and path_parts[0] == 'Icerik':
# Extract the decision ID from the path
decision_id = '/'.join(path_parts[1:]) # e.g., "7288/2021-1303"
return decision_id
except Exception as e:
logger.debug(f"Could not extract decision ID from URL {url}: {e}")
return None
def _extract_decision_metadata_from_title(self, title: str) -> Dict[str, Optional[str]]:
"""Extract decision metadata from title string."""
metadata = {
"decision_date": None,
"decision_number": None
}
if not title:
return metadata
# Extract decision date (DD/MM/YYYY format)
date_match = re.search(r'(\d{1,2}/\d{1,2}/\d{4})', title)
if date_match:
metadata["decision_date"] = date_match.group(1)
# Extract decision number (YYYY/XXXX format)
number_match = re.search(r'(\d{4}/\d+)', title)
if number_match:
metadata["decision_number"] = number_match.group(1)
return metadata
async def search_decisions(self, params: KvkkSearchRequest) -> KvkkSearchResult:
"""Search for KVKK decisions using Brave API."""
search_query = self._construct_search_query(params.keywords)
logger.info(f"KvkkApiClient: Searching with query: {search_query}")
try:
# Calculate offset for pagination
offset = (params.page - 1) * params.pageSize
response = await self.http_client.get(
self.BRAVE_API_URL,
headers={
"Accept": "application/json",
"Accept-Encoding": "gzip",
"x-subscription-token": self.brave_api_token
},
params={
"q": search_query,
"country": "TR",
"search_lang": "tr",
"ui_lang": "tr-TR",
"offset": offset,
"count": params.pageSize
}
)
response.raise_for_status()
data = response.json()
# Extract search results
decisions = []
web_results = data.get("web", {}).get("results", [])
for result in web_results:
title = result.get("title", "")
url = result.get("url", "")
description = result.get("description", "")
# Extract metadata from title
metadata = self._extract_decision_metadata_from_title(title)
# Extract decision ID from URL
decision_id = self._extract_decision_id_from_url(url)
decision = KvkkDecisionSummary(
title=title,
url=HttpUrl(url) if url else None,
description=description,
decision_id=decision_id,
publication_date=metadata.get("decision_date"),
decision_number=metadata.get("decision_number")
)
decisions.append(decision)
# Get total results if available
total_results = None
query_info = data.get("query", {})
if "total_results" in query_info:
total_results = query_info["total_results"]
return KvkkSearchResult(
decisions=decisions,
total_results=total_results,
page=params.page,
pageSize=params.pageSize,
query=search_query
)
except httpx.RequestError as e:
logger.error(f"KvkkApiClient: HTTP request error during search: {e}")
return KvkkSearchResult(
decisions=[],
total_results=0,
page=params.page,
pageSize=params.pageSize,
query=search_query
)
except Exception as e:
logger.error(f"KvkkApiClient: Unexpected error during search: {e}")
return KvkkSearchResult(
decisions=[],
total_results=0,
page=params.page,
pageSize=params.pageSize,
query=search_query
)
def _extract_decision_content_from_html(self, html: str, url: str) -> Dict[str, Any]:
"""Extract decision content from KVKK decision page HTML."""
try:
soup = BeautifulSoup(html, 'html.parser')
# Extract title
title = None
title_element = soup.find('h3', class_='blog-post-title')
if title_element:
title = title_element.get_text(strip=True)
elif soup.title:
title = soup.title.get_text(strip=True)
# Extract decision content from the main content div
content_div = soup.find('div', class_='blog-post-inner')
if not content_div:
# Fallback to other possible content containers
content_div = soup.find('div', style='text-align:justify;')
if not content_div:
logger.warning(f"Could not find decision content div in {url}")
return {
"title": title,
"decision_date": None,
"decision_number": None,
"subject_summary": None,
"html_content": None
}
# Extract decision metadata from table
decision_date = None
decision_number = None
subject_summary = None
table = content_div.find('table')
if table:
rows = table.find_all('tr')
for row in rows:
cells = row.find_all('td')
if len(cells) >= 3:
field_name = cells[0].get_text(strip=True)
field_value = cells[2].get_text(strip=True)
if 'Karar Tarihi' in field_name:
decision_date = field_value
elif 'Karar No' in field_name:
decision_number = field_value
elif 'Konu Özeti' in field_name:
subject_summary = field_value
return {
"title": title,
"decision_date": decision_date,
"decision_number": decision_number,
"subject_summary": subject_summary,
"html_content": str(content_div)
}
except Exception as e:
logger.error(f"Error extracting content from HTML for {url}: {e}")
return {
"title": None,
"decision_date": None,
"decision_number": None,
"subject_summary": None,
"html_content": None
}
def _convert_html_to_markdown(self, html_content: str) -> Optional[str]:
"""Convert HTML content to Markdown using MarkItDown with BytesIO to avoid filename length issues."""
if not html_content:
return None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_content.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown(enable_plugins=False)
result = md_converter.convert(html_stream)
return result.text_content
except Exception as e:
logger.error(f"Error converting HTML to Markdown: {e}")
return None
async def get_decision_document(self, decision_url: str, page_number: int = 1) -> KvkkDocumentMarkdown:
"""Retrieve and convert a KVKK decision document to paginated Markdown."""
logger.info(f"KvkkApiClient: Getting decision document from: {decision_url}, page: {page_number}")
try:
# Fetch the decision page
response = await self.http_client.get(decision_url)
response.raise_for_status()
# Extract content from HTML
extracted_data = self._extract_decision_content_from_html(response.text, decision_url)
# Convert HTML content to Markdown
full_markdown_content = None
if extracted_data["html_content"]:
full_markdown_content = self._convert_html_to_markdown(extracted_data["html_content"])
if not full_markdown_content:
return KvkkDocumentMarkdown(
source_url=HttpUrl(decision_url),
title=extracted_data["title"],
decision_date=extracted_data["decision_date"],
decision_number=extracted_data["decision_number"],
subject_summary=extracted_data["subject_summary"],
markdown_chunk=None,
current_page=page_number,
total_pages=0,
is_paginated=False,
error_message="Could not convert document content to Markdown"
)
# Calculate pagination
content_length = len(full_markdown_content)
total_pages = math.ceil(content_length / self.DOCUMENT_MARKDOWN_CHUNK_SIZE)
if total_pages == 0:
total_pages = 1
# Clamp page number to valid range
current_page_clamped = max(1, min(page_number, total_pages))
# Extract the requested chunk
start_index = (current_page_clamped - 1) * self.DOCUMENT_MARKDOWN_CHUNK_SIZE
end_index = start_index + self.DOCUMENT_MARKDOWN_CHUNK_SIZE
markdown_chunk = full_markdown_content[start_index:end_index]
return KvkkDocumentMarkdown(
source_url=HttpUrl(decision_url),
title=extracted_data["title"],
decision_date=extracted_data["decision_date"],
decision_number=extracted_data["decision_number"],
subject_summary=extracted_data["subject_summary"],
markdown_chunk=markdown_chunk,
current_page=current_page_clamped,
total_pages=total_pages,
is_paginated=(total_pages > 1),
error_message=None
)
except httpx.HTTPStatusError as e:
error_msg = f"HTTP error {e.response.status_code} when fetching decision document"
logger.error(f"KvkkApiClient: {error_msg}")
return KvkkDocumentMarkdown(
source_url=HttpUrl(decision_url),
title=None,
decision_date=None,
decision_number=None,
subject_summary=None,
markdown_chunk=None,
current_page=page_number,
total_pages=0,
is_paginated=False,
error_message=error_msg
)
except Exception as e:
error_msg = f"Unexpected error when fetching decision document: {str(e)}"
logger.error(f"KvkkApiClient: {error_msg}")
return KvkkDocumentMarkdown(
source_url=HttpUrl(decision_url),
title=None,
decision_date=None,
decision_number=None,
subject_summary=None,
markdown_chunk=None,
current_page=page_number,
total_pages=0,
is_paginated=False,
error_message=error_msg
)
async def close_client_session(self):
"""Close the HTTP client session."""
if hasattr(self, 'http_client') and self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("KvkkApiClient: HTTP client session closed.")
@@ -1,49 +0,0 @@
# kvkk_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional, Any
class KvkkSearchRequest(BaseModel):
"""Model for KVKK (Personal Data Protection Authority) search request via Brave API."""
keywords: str = Field(..., description="""
Keywords to search for in KVKK decisions.
The search will automatically include 'site:kvkk.gov.tr "karar özeti"' to target KVKK decision summaries.
Examples: "açık rıza", "veri güvenliği", "kişisel veri işleme"
""")
page: int = Field(1, ge=1, le=50, description="Page number for search results (1-50).")
pageSize: int = Field(10, ge=1, le=10, description="Number of results per page (1-10).")
class KvkkDecisionSummary(BaseModel):
"""Model for a single KVKK decision summary from Brave search results."""
title: Optional[str] = Field(None, description="Decision title from search results.")
url: Optional[HttpUrl] = Field(None, description="URL to the KVKK decision page.")
description: Optional[str] = Field(None, description="Brief description or snippet from search results.")
decision_id: Optional[str] = Field(None, description="Value")
publication_date: Optional[str] = Field(None, description="Value")
decision_number: Optional[str] = Field(None, description="Value")
class KvkkSearchResult(BaseModel):
"""Model for the overall search result for KVKK decisions."""
decisions: List[KvkkDecisionSummary] = Field(default_factory=list, description="List of KVKK decisions found.")
total_results: Optional[int] = Field(None, description="Value")
page: int = Field(1, description="Current page number of results.")
pageSize: int = Field(10, description="Number of results per page.")
query: Optional[str] = Field(None, description="The actual search query sent to Brave API.")
class KvkkDocumentMarkdown(BaseModel):
"""Model for KVKK decision document content converted to paginated Markdown."""
source_url: HttpUrl = Field(description="URL of the original KVKK decision page.")
title: Optional[str] = Field(None, description="Title of the KVKK decision.")
decision_date: Optional[str] = Field(None, description="Decision date (Karar Tarihi).")
decision_number: Optional[str] = Field(None, description="Decision number (Karar No).")
subject_summary: Optional[str] = Field(None, description="Subject summary (Konu Özeti).")
markdown_chunk: Optional[str] = Field(None, description="A 5,000 character chunk of the Markdown content.")
current_page: int = Field(description="The current page number of the markdown chunk (1-indexed).")
total_pages: int = Field(description="Total number of pages for the full markdown content.")
is_paginated: bool = Field(description="True if the full markdown content is split into multiple pages.")
error_message: Optional[str] = Field(None, description="Value")
class Config:
json_encoders = {
HttpUrl: str
}
@@ -1,28 +0,0 @@
"""
MCP Auth Toolkit - OAuth 2.1 + Authorization for Model Context Protocol Servers
Integrated with Clerk Authentication
"""
from .middleware import (
AuthContext,
FastMCPAuthWrapper,
MCPAuthMiddleware,
auth_required,
)
from .oauth import OAuthConfig, OAuthProvider
from .policy import PolicyEngine, ToolPolicy, create_default_policies
from .storage import PersistentStorage
__version__ = "0.1.0"
__all__ = [
"OAuthProvider",
"OAuthConfig",
"AuthContext",
"auth_required",
"create_default_policies",
"MCPAuthMiddleware",
"FastMCPAuthWrapper",
"PolicyEngine",
"ToolPolicy",
"PersistentStorage",
]
@@ -1,73 +0,0 @@
"""
Clerk OAuth configuration for MCP Auth Toolkit
"""
import os
import logging
from .oauth import OAuthConfig
logger = logging.getLogger(__name__)
def create_clerk_oauth_config() -> OAuthConfig:
"""Create OAuth configuration for Clerk integration using SDK"""
# Get Clerk configuration from environment
clerk_domain = os.getenv("CLERK_DOMAIN", "accounts.yargimcp.com")
clerk_publishable_key = os.getenv("CLERK_PUBLISHABLE_KEY")
clerk_secret_key = os.getenv("CLERK_SECRET_KEY")
if not clerk_publishable_key or not clerk_secret_key:
raise ValueError("CLERK_PUBLISHABLE_KEY and CLERK_SECRET_KEY are required")
# For Clerk with custom domains, we use our adapter endpoints
# This allows us to handle the custom domain flow properly
base_url = os.getenv("BASE_URL", "https://yargimcp.com")
config = OAuthConfig(
client_id=clerk_publishable_key,
client_secret=clerk_secret_key,
# Use our adapter endpoints instead of Clerk's direct endpoints
authorization_endpoint=f"{base_url}/authorize",
token_endpoint=f"{base_url}/token",
# Keep Clerk's JWKS for token validation
jwks_uri=f"https://{clerk_domain}/.well-known/jwks.json",
issuer=base_url, # We're the issuer for MCP tokens
scopes=["mcp:tools:read", "mcp:tools:write", "openid", "profile", "email"]
)
logger.info(f"Created Clerk OAuth config with adapter endpoints")
logger.info(f"Clerk domain: {clerk_domain}")
logger.debug(f"Authorization endpoint: {config.authorization_endpoint}")
logger.debug(f"Token endpoint: {config.token_endpoint}")
return config
def get_jwt_secret() -> str:
"""Get JWT secret for token signing"""
jwt_secret = os.getenv("JWT_SECRET_KEY")
if not jwt_secret:
raise ValueError("JWT_SECRET_KEY environment variable is required")
return jwt_secret
def create_mcp_server_config():
"""Create complete MCP server configuration for Clerk integration"""
try:
oauth_config = create_clerk_oauth_config()
jwt_secret = get_jwt_secret()
return {
"oauth_config": oauth_config,
"jwt_secret": jwt_secret,
"base_url": os.getenv("BASE_URL", "https://yargi-mcp.fly.dev"),
"auth_enabled": os.getenv("ENABLE_AUTH", "true").lower() == "true"
}
except Exception as e:
logger.error(f"Failed to create MCP server config: {e}")
raise
@@ -1,315 +0,0 @@
"""
MCP server middleware for OAuth authentication and authorization
"""
import functools
import logging
from collections.abc import Callable
from dataclasses import dataclass
from typing import Any, Optional
logger = logging.getLogger(__name__)
try:
from fastmcp import FastMCP
FASTMCP_AVAILABLE = True
except ImportError:
FASTMCP_AVAILABLE = False
FastMCP = None
logger.warning("FastMCP not available, some features will be disabled")
from .oauth import OAuthProvider
from .policy import PolicyEngine
@dataclass
class AuthContext:
"""Authentication context passed to MCP tools"""
user_id: str
scopes: list[str]
claims: dict[str, Any]
token: str
class MCPAuthMiddleware:
"""Authentication middleware for MCP servers"""
def __init__(self, oauth_provider: OAuthProvider, policy_engine: PolicyEngine):
self.oauth_provider = oauth_provider
self.policy_engine = policy_engine
def authenticate_request(self, authorization_header: str) -> AuthContext | None:
"""Extract and validate auth token from request"""
if not authorization_header:
logger.debug("No authorization header provided")
return None
if not authorization_header.startswith("Bearer "):
logger.debug("Authorization header does not start with 'Bearer '")
return None
token = authorization_header[7:] # Remove 'Bearer ' prefix
token_info = self.oauth_provider.introspect_token(token)
if not token_info.get("active"):
logger.warning("Token is not active")
return None
logger.debug(f"Authenticated user: {token_info.get('sub', 'unknown')}")
return AuthContext(
user_id=token_info.get("sub", "unknown"),
scopes=token_info.get("mcp_tool_scopes", []),
claims=token_info,
token=token,
)
def authorize_tool_call(
self, tool_name: str, auth_context: AuthContext
) -> tuple[bool, str | None]:
"""Check if user can call the specified tool"""
return self.policy_engine.authorize_tool_call(
tool_name=tool_name,
user_scopes=auth_context.scopes,
user_claims=auth_context.claims,
)
def auth_required(
oauth_provider: OAuthProvider,
policy_engine: PolicyEngine,
tool_name: str | None = None,
):
"""
Decorator to require authentication for MCP tool functions
Usage:
@auth_required(oauth_provider, policy_engine, "search_yargitay")
def my_tool_function(context: AuthContext, ...):
pass
"""
def decorator(func: Callable) -> Callable:
middleware = MCPAuthMiddleware(oauth_provider, policy_engine)
@functools.wraps(func)
async def wrapper(*args, **kwargs):
# Extract authorization header from kwargs
auth_header = kwargs.pop("authorization", None)
# Also check in args if it's a Request object
if not auth_header and args:
for arg in args:
if hasattr(arg, 'headers'):
auth_header = arg.headers.get("Authorization")
break
if not auth_header:
logger.warning(f"No authorization header for tool '{tool_name or func.__name__}'")
raise PermissionError("Authorization header required")
auth_context = middleware.authenticate_request(auth_header)
if not auth_context:
logger.warning(f"Authentication failed for tool '{tool_name or func.__name__}'")
raise PermissionError("Invalid or expired token")
actual_tool_name = tool_name or func.__name__
authorized, reason = middleware.authorize_tool_call(
actual_tool_name, auth_context
)
if not authorized:
logger.warning(f"Authorization failed for tool '{actual_tool_name}': {reason}")
raise PermissionError(f"Access denied: {reason}")
# Add auth context to function call
return await func(auth_context, *args, **kwargs)
return wrapper
return decorator
class FastMCPAuthWrapper:
"""Wrapper for FastMCP servers to add authentication"""
def __init__(
self,
mcp_server: "FastMCP",
oauth_provider: OAuthProvider,
policy_engine: PolicyEngine,
):
if not FASTMCP_AVAILABLE:
raise ImportError("FastMCP is required for FastMCPAuthWrapper")
self.mcp_server = mcp_server
self.middleware = MCPAuthMiddleware(oauth_provider, policy_engine)
self.oauth_provider = oauth_provider
logger.info("Initializing FastMCP authentication wrapper")
self._wrap_tools()
def _wrap_tools(self):
"""Wrap all existing tools with auth middleware"""
# Try different FastMCP tool storage locations
tool_registry = None
if hasattr(self.mcp_server, '_tools'):
tool_registry = self.mcp_server._tools
elif hasattr(self.mcp_server, 'tools'):
tool_registry = self.mcp_server.tools
elif hasattr(self.mcp_server, '_tool_registry'):
tool_registry = self.mcp_server._tool_registry
elif hasattr(self.mcp_server, '_handlers') and hasattr(self.mcp_server._handlers, 'tools'):
tool_registry = self.mcp_server._handlers.tools
if not tool_registry:
logger.warning("FastMCP server tool registry not found, tools will not be automatically wrapped")
logger.debug(f"Available server attributes: {dir(self.mcp_server)}")
return
logger.debug(f"Found tool registry with {len(tool_registry)} tools")
original_tools = dict(tool_registry)
wrapped_count = 0
for tool_name, tool_func in original_tools.items():
try:
wrapped_func = self._create_auth_wrapper(tool_name, tool_func)
tool_registry[tool_name] = wrapped_func
wrapped_count += 1
logger.debug(f"Wrapped tool: {tool_name}")
except Exception as e:
logger.error(f"Failed to wrap tool {tool_name}: {e}")
logger.info(f"Successfully wrapped {wrapped_count} tools with authentication")
def _create_auth_wrapper(self, tool_name: str, original_func: Callable) -> Callable:
"""Create auth wrapper for a specific tool"""
@functools.wraps(original_func)
async def auth_wrapper(*args, **kwargs):
# Extract authorization from various sources
auth_header = None
# Check kwargs first
auth_header = kwargs.pop("authorization", None)
# Check if first argument is a Request object
if not auth_header and args:
first_arg = args[0]
if hasattr(first_arg, 'headers'):
auth_header = first_arg.headers.get("Authorization")
if not auth_header:
logger.warning(f"No authorization header for tool '{tool_name}'")
raise PermissionError("Authorization required")
auth_context = self.middleware.authenticate_request(auth_header)
if not auth_context:
logger.warning(f"Authentication failed for tool '{tool_name}'")
raise PermissionError("Invalid token")
authorized, reason = self.middleware.authorize_tool_call(
tool_name, auth_context
)
if not authorized:
logger.warning(f"Authorization failed for tool '{tool_name}': {reason}")
raise PermissionError(f"Access denied: {reason}")
# Add auth context to kwargs
kwargs["auth_context"] = auth_context
logger.debug(f"Calling tool '{tool_name}' for user {auth_context.user_id}")
return await original_func(*args, **kwargs)
return auth_wrapper
def add_oauth_endpoints(self):
"""Add OAuth endpoints to the MCP server"""
@self.mcp_server.tool(
description="Initiate OAuth 2.1 authorization flow with PKCE",
annotations={"readOnlyHint": True, "idempotentHint": False}
)
async def oauth_authorize(redirect_uri: str, scopes: Optional[str] = None):
"""OAuth authorization endpoint"""
scope_list = scopes.split(" ") if scopes else None
auth_url, pkce = self.oauth_provider.generate_authorization_url(
redirect_uri=redirect_uri, scopes=scope_list
)
logger.info(f"Generated authorization URL for redirect_uri: {redirect_uri}")
return {
"authorization_url": auth_url,
"code_verifier": pkce.verifier, # For PKCE flow
"code_challenge": pkce.challenge,
"instructions": "Use the authorization_url to complete OAuth flow, then exchange the returned code using oauth_token tool"
}
@self.mcp_server.tool(
description="Exchange OAuth authorization code for access token",
annotations={"readOnlyHint": False, "idempotentHint": False}
)
async def oauth_token(
code: str,
state: str,
redirect_uri: str
):
"""OAuth token exchange endpoint"""
try:
result = await self.oauth_provider.exchange_code_for_token(
code=code, state=state, redirect_uri=redirect_uri
)
logger.info("Successfully exchanged authorization code for token")
return result
except Exception as e:
logger.error(f"Token exchange failed: {e}")
raise
@self.mcp_server.tool(
description="Validate and introspect OAuth access token",
annotations={"readOnlyHint": True, "idempotentHint": True}
)
async def oauth_introspect(token: str):
"""Token introspection endpoint"""
result = self.oauth_provider.introspect_token(token)
logger.debug(f"Token introspection: active={result.get('active', False)}")
return result
@self.mcp_server.tool(
description="Revoke OAuth access token",
annotations={"readOnlyHint": False, "idempotentHint": False}
)
async def oauth_revoke(token: str):
"""Token revocation endpoint"""
success = self.oauth_provider.revoke_token(token)
logger.info(f"Token revocation: success={success}")
return {"revoked": success}
@self.mcp_server.tool(
description="Get list of tools available to authenticated user",
annotations={"readOnlyHint": True, "idempotentHint": True}
)
async def oauth_user_tools(authorization: str):
"""Get user's allowed tools based on scopes"""
auth_context = self.middleware.authenticate_request(authorization)
if not auth_context:
raise PermissionError("Invalid token")
allowed_patterns = self.middleware.policy_engine.get_allowed_tools(auth_context.scopes)
return {
"user_id": auth_context.user_id,
"scopes": auth_context.scopes,
"allowed_tool_patterns": allowed_patterns,
"message": "Use these patterns to determine which tools you can access"
}
logger.info("Added OAuth endpoints: oauth_authorize, oauth_token, oauth_introspect, oauth_revoke, oauth_user_tools")
@@ -1,304 +0,0 @@
"""
OAuth 2.1 + PKCE implementation for MCP servers with Clerk integration
"""
import base64
import hashlib
import secrets
import time
import logging
from dataclasses import dataclass
from datetime import datetime, timedelta
from typing import Any, Optional
from urllib.parse import urlencode
import httpx
import jwt
from jwt.exceptions import PyJWTError, InvalidTokenError
from .storage import PersistentStorage
# Try to import Clerk SDK
try:
from clerk_backend_api import Clerk
CLERK_AVAILABLE = True
except ImportError:
CLERK_AVAILABLE = False
Clerk = None
logger = logging.getLogger(__name__)
@dataclass
class OAuthConfig:
"""OAuth provider configuration for Clerk"""
client_id: str
client_secret: str
authorization_endpoint: str
token_endpoint: str
jwks_uri: str | None = None
issuer: str = "mcp-auth"
scopes: list[str] = None
def __post_init__(self):
if self.scopes is None:
self.scopes = ["mcp:tools:read", "mcp:tools:write"]
class PKCEChallenge:
"""PKCE challenge/verifier pair for OAuth 2.1"""
def __init__(self):
self.verifier = (
base64.urlsafe_b64encode(secrets.token_bytes(32))
.decode("utf-8")
.rstrip("=")
)
challenge_bytes = hashlib.sha256(self.verifier.encode("utf-8")).digest()
self.challenge = (
base64.urlsafe_b64encode(challenge_bytes).decode("utf-8").rstrip("=")
)
class OAuthProvider:
"""OAuth 2.1 provider with PKCE support and Clerk integration"""
def __init__(self, config: OAuthConfig, jwt_secret: str):
self.config = config
self.jwt_secret = jwt_secret
# Use persistent storage instead of memory
self.storage = PersistentStorage()
# Initialize Clerk SDK if available
self.clerk = None
if CLERK_AVAILABLE and config.client_secret:
try:
self.clerk = Clerk(bearer_auth=config.client_secret)
logger.info("Clerk SDK initialized successfully")
except Exception as e:
logger.warning(f"Failed to initialize Clerk SDK: {e}")
logger.info("OAuth provider initialized with persistent storage")
def generate_authorization_url(
self,
redirect_uri: str,
state: str | None = None,
scopes: list[str] | None = None,
) -> tuple[str, PKCEChallenge]:
"""Generate OAuth authorization URL with PKCE for Clerk"""
pkce = PKCEChallenge()
session_id = secrets.token_urlsafe(32)
if state is None:
state = secrets.token_urlsafe(16)
if scopes is None:
scopes = self.config.scopes
# Store session data with expiration
session_data = {
"pkce_verifier": pkce.verifier,
"state": state,
"redirect_uri": redirect_uri,
"scopes": scopes,
"created_at": time.time(),
"expires_at": (datetime.utcnow() + timedelta(minutes=10)).timestamp(),
}
self.storage.set_session(session_id, session_data)
# Build Clerk OAuth URL
# Check if this is a custom domain (sign-in endpoint)
if self.config.authorization_endpoint.endswith('/sign-in'):
# For custom domains, Clerk expects redirect_url parameter
params = {
"redirect_url": redirect_uri,
"state": f"{state}:{session_id}",
}
auth_url = f"{self.config.authorization_endpoint}?{urlencode(params)}"
else:
# Standard OAuth flow with PKCE
params = {
"response_type": "code",
"client_id": self.config.client_id,
"redirect_uri": redirect_uri,
"scope": " ".join(scopes),
"state": f"{state}:{session_id}", # Combine state with session ID
"code_challenge": pkce.challenge,
"code_challenge_method": "S256",
}
auth_url = f"{self.config.authorization_endpoint}?{urlencode(params)}"
logger.info(f"Generated OAuth URL with session {session_id[:8]}...")
logger.debug(f"Auth URL: {auth_url}")
return auth_url, pkce
async def exchange_code_for_token(
self, code: str, state: str, redirect_uri: str
) -> dict[str, Any]:
"""Exchange authorization code for access token with Clerk"""
try:
original_state, session_id = state.split(":", 1)
except ValueError as e:
logger.error(f"Invalid state format: {state}")
raise ValueError("Invalid state format") from e
session = self.storage.get_session(session_id)
if not session:
logger.error(f"Session {session_id} not found")
raise ValueError("Invalid session")
# Check session expiration
if datetime.utcnow().timestamp() > session.get("expires_at", 0):
self.storage.delete_session(session_id)
logger.error(f"Session {session_id} expired")
raise ValueError("Session expired")
if session["state"] != original_state:
logger.error(f"State mismatch: expected {session['state']}, got {original_state}")
raise ValueError("State mismatch")
if session["redirect_uri"] != redirect_uri:
logger.error(f"Redirect URI mismatch: expected {session['redirect_uri']}, got {redirect_uri}")
raise ValueError("Redirect URI mismatch")
# Prepare token exchange request for Clerk
token_data = {
"grant_type": "authorization_code",
"client_id": self.config.client_id,
"client_secret": self.config.client_secret,
"code": code,
"redirect_uri": redirect_uri,
"code_verifier": session["pkce_verifier"],
}
logger.info(f"Exchanging code with Clerk for session {session_id[:8]}...")
async with httpx.AsyncClient() as client:
response = await client.post(
self.config.token_endpoint,
data=token_data,
headers={"Content-Type": "application/x-www-form-urlencoded"},
timeout=30.0,
)
if response.status_code != 200:
logger.error(f"Clerk token exchange failed: {response.status_code} - {response.text}")
raise ValueError(f"Token exchange failed: {response.text}")
token_response = response.json()
logger.info("Successfully exchanged code for Clerk token")
# Create MCP-scoped JWT token
access_token = self._create_mcp_token(
session["scopes"], token_response.get("access_token"), session_id
)
# Store token for introspection
token_id = secrets.token_urlsafe(16)
token_data = {
"access_token": access_token,
"scopes": session["scopes"],
"created_at": time.time(),
"expires_at": (datetime.utcnow() + timedelta(hours=1)).timestamp(),
"session_id": session_id,
"clerk_token": token_response.get("access_token"),
}
self.storage.set_token(token_id, token_data)
# Clean up session
self.storage.delete_session(session_id)
return {
"access_token": access_token,
"token_type": "bearer",
"expires_in": 3600,
"scope": " ".join(session["scopes"]),
}
def validate_pkce(self, code_verifier: str, code_challenge: str) -> bool:
"""Validate PKCE code challenge (RFC 7636)"""
# S256 method
verifier_hash = hashlib.sha256(code_verifier.encode()).digest()
expected_challenge = base64.urlsafe_b64encode(verifier_hash).decode().rstrip('=')
return expected_challenge == code_challenge
def _create_mcp_token(
self, scopes: list[str], upstream_token: str, session_id: str
) -> str:
"""Create MCP-scoped JWT token with Clerk token embedded"""
now = int(time.time())
payload = {
"iss": self.config.issuer,
"sub": session_id,
"aud": "mcp-server",
"iat": now,
"exp": now + 3600, # 1 hour expiration
"mcp_tool_scopes": scopes,
"upstream_token": upstream_token,
"clerk_integration": True,
}
return jwt.encode(payload, self.jwt_secret, algorithm="HS256")
def introspect_token(self, token: str) -> dict[str, Any]:
"""Introspect and validate MCP token"""
try:
payload = jwt.decode(token, self.jwt_secret, algorithms=["HS256"])
# Check if token is expired
if payload.get("exp", 0) < time.time():
return {"active": False, "error": "token_expired"}
return {
"active": True,
"sub": payload.get("sub"),
"aud": payload.get("aud"),
"iss": payload.get("iss"),
"exp": payload.get("exp"),
"iat": payload.get("iat"),
"mcp_tool_scopes": payload.get("mcp_tool_scopes", []),
"upstream_token": payload.get("upstream_token"),
"clerk_integration": payload.get("clerk_integration", False),
}
except PyJWTError as e:
logger.warning(f"Token validation failed: {e}")
return {"active": False, "error": "invalid_token"}
def revoke_token(self, token: str) -> bool:
"""Revoke a token"""
try:
payload = jwt.decode(token, self.jwt_secret, algorithms=["HS256"])
session_id = payload.get("sub")
# Remove all tokens associated with this session
all_tokens = self.storage.get_tokens()
tokens_to_remove = [
token_id
for token_id, token_data in all_tokens.items()
if token_data.get("session_id") == session_id
]
for token_id in tokens_to_remove:
self.storage.delete_token(token_id)
logger.info(f"Revoked {len(tokens_to_remove)} tokens for session {session_id}")
return True
except InvalidTokenError as e:
logger.warning(f"Token revocation failed: {e}")
return False
def cleanup_expired_sessions(self):
"""Clean up expired sessions and tokens"""
# This is now handled automatically by persistent storage
self.storage.cleanup_expired_sessions()
logger.debug("Cleanup completed via persistent storage")
@@ -1,201 +0,0 @@
"""
Authorization policy engine for MCP tools
"""
import re
import logging
from dataclasses import dataclass
from enum import Enum
from typing import Any
logger = logging.getLogger(__name__)
class PolicyAction(Enum):
ALLOW = "allow"
DENY = "deny"
@dataclass
class ToolPolicy:
"""Policy rule for MCP tool access"""
tool_pattern: str # regex pattern for tool names
required_scopes: list[str]
action: PolicyAction = PolicyAction.ALLOW
conditions: dict[str, Any] | None = None
def matches_tool(self, tool_name: str) -> bool:
"""Check if the policy applies to given tool"""
return bool(re.match(self.tool_pattern, tool_name))
def evaluate_scopes(self, user_scopes: list[str]) -> bool:
"""Check if user has required scopes"""
return all(scope in user_scopes for scope in self.required_scopes)
class PolicyEngine:
"""Authorization policy engine for Turkish legal database tools"""
def __init__(self):
self.policies: list[ToolPolicy] = []
self.default_action = PolicyAction.DENY
def add_policy(self, policy: ToolPolicy):
"""Add a policy rule"""
self.policies.append(policy)
logger.debug(f"Added policy: {policy.tool_pattern} -> {policy.required_scopes}")
def add_tool_scope_policy(
self,
tool_pattern: str,
required_scopes: str | list[str],
action: PolicyAction = PolicyAction.ALLOW,
):
"""Convenience method to add tool-scope policy"""
if isinstance(required_scopes, str):
required_scopes = [required_scopes]
policy = ToolPolicy(
tool_pattern=tool_pattern, required_scopes=required_scopes, action=action
)
self.add_policy(policy)
def authorize_tool_call(
self,
tool_name: str,
user_scopes: list[str],
user_claims: dict[str, Any] | None = None,
) -> tuple[bool, str | None]:
"""
Authorize a tool call
Returns:
(authorized: bool, reason: Optional[str])
"""
logger.debug(f"Authorizing tool '{tool_name}' for user with scopes: {user_scopes}")
matching_policies = [
policy for policy in self.policies if policy.matches_tool(tool_name)
]
if not matching_policies:
if self.default_action == PolicyAction.ALLOW:
logger.debug(f"No policies found for '{tool_name}', allowing by default")
return True, None
else:
logger.warning(f"No policies found for '{tool_name}', denying by default")
return False, f"No policy found for tool '{tool_name}', default deny"
# Check for explicit deny policies first
for policy in matching_policies:
if policy.action == PolicyAction.DENY:
if policy.evaluate_scopes(user_scopes):
logger.warning(f"Explicit deny policy matched for '{tool_name}'")
return False, f"Explicit deny policy for tool '{tool_name}'"
# Check allow policies
allow_policies = [
p for p in matching_policies if p.action == PolicyAction.ALLOW
]
if not allow_policies:
logger.warning(f"No allow policies found for '{tool_name}'")
return False, f"No allow policies found for tool '{tool_name}'"
for policy in allow_policies:
if policy.evaluate_scopes(user_scopes):
if self._evaluate_conditions(policy.conditions, user_claims):
logger.debug(f"Authorization granted for '{tool_name}'")
return True, None
logger.warning(f"Insufficient scopes for '{tool_name}'. Required: {[p.required_scopes for p in allow_policies]}, User has: {user_scopes}")
return False, f"Insufficient scopes for tool '{tool_name}'"
def _evaluate_conditions(
self,
conditions: dict[str, Any] | None,
user_claims: dict[str, Any] | None,
) -> bool:
"""Evaluate additional policy conditions"""
if not conditions:
return True
if not user_claims:
logger.debug("No user claims provided, conditions evaluation failed")
return False
for key, expected_value in conditions.items():
user_value = user_claims.get(key)
if isinstance(expected_value, list):
if user_value not in expected_value:
logger.debug(f"Condition failed: {key} = {user_value} not in {expected_value}")
return False
elif user_value != expected_value:
logger.debug(f"Condition failed: {key} = {user_value} != {expected_value}")
return False
return True
def get_allowed_tools(self, user_scopes: list[str]) -> list[str]:
"""Get list of tool patterns user is allowed to call"""
allowed_tools = []
for policy in self.policies:
if policy.action == PolicyAction.ALLOW and policy.evaluate_scopes(
user_scopes
):
allowed_tools.append(policy.tool_pattern)
return allowed_tools
def create_turkish_legal_policies() -> PolicyEngine:
"""Create policy set for Turkish legal database MCP server"""
engine = PolicyEngine()
# Administrative tools (full access)
engine.add_tool_scope_policy(".*", ["mcp:tools:admin"])
# Search tools - require read access
engine.add_tool_scope_policy("search.*", ["mcp:tools:read"])
# Fetch/get document tools - require read access
engine.add_tool_scope_policy("get_.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("fetch.*", ["mcp:tools:read"])
# Specific Turkish legal database tools
engine.add_tool_scope_policy("search_yargitay.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_danistay.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_anayasa.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_rekabet.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_kik.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_emsal.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_uyusmazlik.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_sayistay.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_.*_bedesten", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_yerel_hukuk.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_istinaf_hukuk.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("search_kyb.*", ["mcp:tools:read"])
# Document retrieval tools
engine.add_tool_scope_policy("get_.*_document.*", ["mcp:tools:read"])
engine.add_tool_scope_policy("get_.*_markdown", ["mcp:tools:read"])
# Write operations (if any future tools need them)
engine.add_tool_scope_policy("create_.*", ["mcp:tools:write"])
engine.add_tool_scope_policy("update_.*", ["mcp:tools:write"])
engine.add_tool_scope_policy("delete_.*", ["mcp:tools:write"])
logger.info("Created Turkish legal database policy engine")
return engine
def create_default_policies() -> PolicyEngine:
"""Create a default policy set for MCP servers (backwards compatibility)"""
return create_turkish_legal_policies()
@@ -1,112 +0,0 @@
"""
Persistent storage for OAuth sessions and tokens
"""
import json
import os
import tempfile
import logging
from datetime import datetime
from typing import Dict, Any, Optional
logger = logging.getLogger(__name__)
class PersistentStorage:
"""File-based persistent storage for OAuth data"""
def __init__(self, storage_dir: str = None):
if storage_dir is None:
# Use system temp directory or environment variable
storage_dir = os.environ.get('TEMP', tempfile.gettempdir())
self.storage_dir = os.path.join(storage_dir, 'mcp_oauth_storage')
os.makedirs(self.storage_dir, exist_ok=True)
self.sessions_file = os.path.join(self.storage_dir, 'oauth_sessions.json')
self.tokens_file = os.path.join(self.storage_dir, 'oauth_tokens.json')
logger.info(f"Persistent OAuth storage initialized at: {self.storage_dir}")
def _load_json(self, filepath: str) -> Dict:
"""Load JSON data from file"""
try:
if os.path.exists(filepath):
with open(filepath, 'r', encoding='utf-8') as f:
return json.load(f)
except Exception as e:
logger.error(f"Error loading {filepath}: {e}")
return {}
def _save_json(self, filepath: str, data: Dict):
"""Save JSON data to file"""
try:
with open(filepath, 'w', encoding='utf-8') as f:
json.dump(data, f, indent=2, default=str)
except Exception as e:
logger.error(f"Error saving {filepath}: {e}")
def get_sessions(self) -> Dict[str, Dict[str, Any]]:
"""Get all OAuth sessions"""
data = self._load_json(self.sessions_file)
# Clean expired sessions
now = datetime.utcnow().timestamp()
valid_sessions = {k: v for k, v in data.items()
if v.get('expires_at', 0) > now}
if len(valid_sessions) != len(data):
self._save_json(self.sessions_file, valid_sessions)
return valid_sessions
def set_session(self, session_id: str, data: Dict[str, Any]):
"""Set OAuth session data"""
sessions = self.get_sessions()
sessions[session_id] = data
self._save_json(self.sessions_file, sessions)
def get_session(self, session_id: str) -> Optional[Dict[str, Any]]:
"""Get specific OAuth session data"""
sessions = self.get_sessions()
return sessions.get(session_id)
def delete_session(self, session_id: str):
"""Delete OAuth session"""
sessions = self.get_sessions()
if session_id in sessions:
del sessions[session_id]
self._save_json(self.sessions_file, sessions)
def get_tokens(self) -> Dict[str, Dict[str, Any]]:
"""Get all OAuth tokens"""
data = self._load_json(self.tokens_file)
# Clean expired tokens
now = datetime.utcnow().timestamp()
valid_tokens = {k: v for k, v in data.items()
if v.get('expires_at', 0) > now}
if len(valid_tokens) != len(data):
self._save_json(self.tokens_file, valid_tokens)
return valid_tokens
def set_token(self, token_id: str, token_data: Dict[str, Any]):
"""Set OAuth token data"""
tokens = self.get_tokens()
tokens[token_id] = token_data
self._save_json(self.tokens_file, tokens)
def get_token(self, token_id: str) -> Optional[Dict[str, Any]]:
"""Get specific OAuth token data"""
tokens = self.get_tokens()
return tokens.get(token_id)
def delete_token(self, token_id: str):
"""Delete OAuth token"""
tokens = self.get_tokens()
if token_id in tokens:
del tokens[token_id]
self._save_json(self.tokens_file, tokens)
def cleanup_expired_sessions(self):
"""Clean up expired sessions and tokens"""
# This is handled automatically in get_sessions() and get_tokens()
sessions = self.get_sessions()
tokens = self.get_tokens()
logger.debug(f"Cleanup: {len(sessions)} active sessions, {len(tokens)} active tokens")
@@ -1,193 +0,0 @@
"""
Factory for creating FastMCP app with MCP Auth Toolkit integration
"""
import logging
import os
from typing import Optional
logger = logging.getLogger(__name__)
try:
from fastmcp import FastMCP
FASTMCP_AVAILABLE = True
except ImportError:
FASTMCP_AVAILABLE = False
FastMCP = None
from mcp_auth import (
OAuthProvider,
PolicyEngine,
FastMCPAuthWrapper,
create_default_policies
)
from mcp_auth.clerk_config import create_mcp_server_config
def create_auth_enabled_app(app_name: str = "Yargı MCP Server") -> FastMCP:
"""Create FastMCP app with authentication enabled"""
if not FASTMCP_AVAILABLE:
raise ImportError("FastMCP is required for authenticated MCP server")
logger.info("Creating FastMCP app with MCP Auth Toolkit integration")
# Create base FastMCP app
app = FastMCP(app_name)
# Check if authentication is enabled
auth_enabled = os.getenv("ENABLE_AUTH", "true").lower() == "true"
if not auth_enabled:
logger.info("Authentication disabled, returning basic FastMCP app")
return app
try:
# Get configuration
logger.info("Getting MCP server configuration...")
config = create_mcp_server_config()
logger.info("Configuration loaded successfully")
# Create OAuth provider with Clerk config
logger.info("Creating OAuth provider...")
oauth_provider = OAuthProvider(
config=config["oauth_config"],
jwt_secret=config["jwt_secret"]
)
logger.info("OAuth provider created successfully")
# Create policy engine for Turkish legal database
policy_engine = create_default_policies()
# Store auth components for later wrapping (after tools are defined)
app._oauth_provider = oauth_provider
app._policy_engine = policy_engine
app._auth_config = config
# Add OAuth endpoints immediately
@app.tool(
description="Initiate OAuth 2.1 authorization flow with PKCE",
annotations={"readOnlyHint": True, "idempotentHint": False}
)
async def oauth_authorize(redirect_uri: str, scopes: str = None):
"""OAuth authorization endpoint"""
scope_list = scopes.split(" ") if scopes else ["mcp:tools:read", "mcp:tools:write"]
auth_url, pkce = oauth_provider.generate_authorization_url(
redirect_uri=redirect_uri, scopes=scope_list
)
logger.info(f"Generated authorization URL for redirect_uri: {redirect_uri}")
return {
"authorization_url": auth_url,
"code_verifier": pkce.verifier,
"code_challenge": pkce.challenge,
"instructions": "Use the authorization_url to complete OAuth flow, then exchange the returned code using oauth_token tool"
}
@app.tool(
description="Exchange OAuth authorization code for access token",
annotations={"readOnlyHint": False, "idempotentHint": False}
)
async def oauth_token(code: str, state: str, redirect_uri: str):
"""OAuth token exchange endpoint"""
try:
result = await oauth_provider.exchange_code_for_token(
code=code, state=state, redirect_uri=redirect_uri
)
logger.info("Successfully exchanged authorization code for token")
return result
except Exception as e:
logger.error(f"Token exchange failed: {e}")
raise
@app.tool(
description="Validate and introspect OAuth access token",
annotations={"readOnlyHint": True, "idempotentHint": True}
)
async def oauth_introspect(token: str):
"""Token introspection endpoint"""
result = oauth_provider.introspect_token(token)
logger.debug(f"Token introspection: active={result.get('active', False)}")
return result
@app.tool(
description="Revoke OAuth access token",
annotations={"readOnlyHint": False, "idempotentHint": False}
)
async def oauth_revoke(token: str):
"""Token revocation endpoint"""
success = oauth_provider.revoke_token(token)
logger.info(f"Token revocation: success={success}")
return {"revoked": success}
logger.info("Successfully created authenticated FastMCP app")
except Exception as e:
logger.error(f"Failed to create authenticated app: {e}")
logger.info("Falling back to non-authenticated FastMCP app")
# Return basic app if auth setup fails
return app
return app
def create_app() -> FastMCP:
"""Create FastMCP app (backwards compatible with mcp_factory.py)"""
return create_auth_enabled_app()
def get_auth_wrapper(app: FastMCP) -> Optional[FastMCPAuthWrapper]:
"""Get auth wrapper from app if available"""
return getattr(app, '_auth_wrapper', None)
def get_oauth_provider(app: FastMCP) -> Optional[OAuthProvider]:
"""Get OAuth provider from app if available"""
return getattr(app, '_oauth_provider', None)
def get_policy_engine(app: FastMCP) -> Optional[PolicyEngine]:
"""Get policy engine from app if available"""
return getattr(app, '_policy_engine', None)
def is_auth_enabled(app: FastMCP) -> bool:
"""Check if authentication is enabled for the app"""
return hasattr(app, '_oauth_provider') or hasattr(app, '_auth_wrapper')
def enable_tool_authentication(app: FastMCP):
"""Enable authentication on all existing tools (call after tools are defined)"""
if not is_auth_enabled(app):
logger.debug("Authentication not enabled, skipping tool authentication")
return
oauth_provider = get_oauth_provider(app)
policy_engine = get_policy_engine(app)
if not oauth_provider or not policy_engine:
logger.warning("OAuth provider or policy engine not available")
return
try:
# Create auth wrapper and wrap tools
auth_wrapper = FastMCPAuthWrapper(
mcp_server=app,
oauth_provider=oauth_provider,
policy_engine=policy_engine
)
# Store wrapper for reference
app._auth_wrapper = auth_wrapper
logger.info("Tool authentication enabled successfully")
except Exception as e:
logger.error(f"Failed to enable tool authentication: {e}")
def cleanup_auth_sessions(app: FastMCP):
"""Clean up expired auth sessions and tokens"""
oauth_provider = get_oauth_provider(app)
if oauth_provider:
oauth_provider.cleanup_expired_sessions()
logger.debug("Cleaned up expired OAuth sessions")
@@ -1,383 +0,0 @@
"""
HTTP adapter for MCP Auth Toolkit OAuth endpoints
Exposes MCP OAuth tools as HTTP endpoints for Claude.ai integration
"""
import os
import logging
import secrets
import time
from typing import Optional
from urllib.parse import urlencode, quote
from datetime import datetime, timedelta
from fastapi import APIRouter, Request, Query, HTTPException
from fastapi.responses import RedirectResponse, JSONResponse
# Try to import Clerk SDK
try:
from clerk_backend_api import Clerk
CLERK_AVAILABLE = True
except ImportError as e:
CLERK_AVAILABLE = False
Clerk = None
logger = logging.getLogger(__name__)
router = APIRouter()
# OAuth configuration
BASE_URL = os.getenv("BASE_URL", "https://yargimcp.com")
@router.get("/.well-known/oauth-authorization-server")
async def get_oauth_metadata():
"""OAuth 2.0 Authorization Server Metadata (RFC 8414)"""
return JSONResponse({
"issuer": BASE_URL,
"authorization_endpoint": f"{BASE_URL}/authorize",
"token_endpoint": f"{BASE_URL}/token",
"registration_endpoint": f"{BASE_URL}/register",
"response_types_supported": ["code"],
"grant_types_supported": ["authorization_code", "refresh_token"],
"code_challenge_methods_supported": ["S256"],
"token_endpoint_auth_methods_supported": ["none"],
"scopes_supported": ["mcp:tools:read", "mcp:tools:write", "openid", "profile", "email"],
"service_documentation": f"{BASE_URL}/mcp/"
})
@router.get("/.well-known/oauth-protected-resource")
async def get_protected_resource_metadata():
"""OAuth Protected Resource Metadata (RFC 9728)"""
return JSONResponse({
"resource": BASE_URL,
"authorization_servers": [BASE_URL],
"bearer_methods_supported": ["header"],
"scopes_supported": ["mcp:tools:read", "mcp:tools:write"],
"resource_documentation": f"{BASE_URL}/docs"
})
@router.get("/authorize")
async def authorize_endpoint(
response_type: str = Query(...),
client_id: str = Query(...),
redirect_uri: str = Query(...),
code_challenge: str = Query(...),
code_challenge_method: str = Query("S256"),
state: Optional[str] = Query(None),
scope: Optional[str] = Query(None)
):
"""OAuth 2.1 Authorization Endpoint - Uses Clerk SDK for custom domains"""
logger.info(f"OAuth authorize request - client_id: {client_id}, redirect_uri: {redirect_uri}")
if not CLERK_AVAILABLE:
logger.error("Clerk SDK not available")
raise HTTPException(status_code=500, detail="Clerk SDK not available")
# Store OAuth session for later validation
try:
from mcp_server_main import app as mcp_app
from mcp_auth_factory import get_oauth_provider
oauth_provider = get_oauth_provider(mcp_app)
if not oauth_provider:
raise HTTPException(status_code=500, detail="OAuth provider not configured")
# Generate session and store PKCE
session_id = secrets.token_urlsafe(32)
if state is None:
state = secrets.token_urlsafe(16)
# Create PKCE challenge
from mcp_auth.oauth import PKCEChallenge
pkce = PKCEChallenge()
# Store session data
session_data = {
"pkce_verifier": pkce.verifier,
"pkce_challenge": code_challenge, # Store the client's challenge
"state": state,
"redirect_uri": redirect_uri,
"client_id": client_id,
"scopes": scope.split(" ") if scope else ["mcp:tools:read", "mcp:tools:write"],
"created_at": time.time(),
"expires_at": (datetime.utcnow() + timedelta(minutes=10)).timestamp(),
}
oauth_provider.storage.set_session(session_id, session_data)
# For Clerk with custom domains, we need to use their hosted sign-in page
# We'll pass our callback URL and session info in the state
callback_url = f"{BASE_URL}/auth/callback"
# Encode session info in state for retrieval after Clerk auth
combined_state = f"{state}:{session_id}"
# Use Clerk's sign-in URL with proper parameters
clerk_domain = os.getenv("CLERK_DOMAIN", "accounts.yargimcp.com")
sign_in_params = {
"redirect_url": f"{callback_url}?state={quote(combined_state)}",
}
sign_in_url = f"https://{clerk_domain}/sign-in?{urlencode(sign_in_params)}"
logger.info(f"Redirecting to Clerk sign-in: {sign_in_url}")
return RedirectResponse(url=sign_in_url)
except Exception as e:
logger.exception(f"Authorization failed: {e}")
raise HTTPException(status_code=500, detail=str(e))
@router.get("/auth/callback")
async def oauth_callback(
request: Request,
state: Optional[str] = Query(None),
clerk_token: Optional[str] = Query(None)
):
"""Handle OAuth callback from Clerk - supports both JWT token and cookie auth"""
logger.info(f"OAuth callback received - state: {state}")
logger.info(f"Query params: {dict(request.query_params)}")
logger.info(f"Cookies: {dict(request.cookies)}")
logger.info(f"Clerk JWT token provided: {bool(clerk_token)}")
# Support both JWT token (for cross-domain) and cookie auth (for subdomain)
try:
if not state:
logger.error("No state parameter provided")
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "Missing state parameter"}
)
# Parse state to get original state and session ID
try:
if ":" in state:
original_state, session_id = state.rsplit(":", 1)
else:
original_state = state
session_id = state # Fallback
except ValueError:
logger.error(f"Invalid state format: {state}")
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "Invalid state format"}
)
# Get OAuth provider
from mcp_server_main import app as mcp_app
from mcp_auth_factory import get_oauth_provider
oauth_provider = get_oauth_provider(mcp_app)
if not oauth_provider:
raise HTTPException(status_code=500, detail="OAuth provider not configured")
# Get stored session
oauth_session = oauth_provider.storage.get_session(session_id)
if not oauth_session:
logger.error(f"OAuth session not found for ID: {session_id}")
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "OAuth session expired or not found"}
)
# Check if we have a JWT token (for cross-domain auth)
user_authenticated = False
auth_method = "none"
if clerk_token:
logger.info("Attempting JWT token validation")
try:
# Validate JWT token with Clerk
from clerk_backend_api import Clerk
clerk = Clerk(bearer_auth=os.getenv("CLERK_SECRET_KEY"))
# Extract session_id from JWT token and verify with Clerk
import jwt
decoded_token = jwt.decode(clerk_token, options={"verify_signature": False})
session_id = decoded_token.get("sid") or decoded_token.get("session_id")
if session_id:
# Verify with Clerk using session_id
session = clerk.sessions.verify(session_id=session_id, token=clerk_token)
user_id = session.user_id if session else None
else:
user_id = None
if user_id:
logger.info(f"JWT token validation successful - user_id: {user_id}")
user_authenticated = True
auth_method = "jwt_token"
# Store user info in session for token exchange
oauth_session["user_id"] = user_id
oauth_session["auth_method"] = "jwt_token"
else:
logger.error("JWT token validation failed - no user_id in claims")
except Exception as e:
logger.error(f"JWT token validation failed: {str(e)}")
# Fall through to cookie validation
# If no JWT token or validation failed, check cookies
if not user_authenticated:
logger.info("Checking for Clerk session cookies")
# Check for Clerk session cookies (for subdomain auth)
clerk_session_cookie = request.cookies.get("__session")
if clerk_session_cookie:
logger.info("Found Clerk session cookie, assuming authenticated")
user_authenticated = True
auth_method = "cookie"
oauth_session["auth_method"] = "cookie"
else:
logger.info("No Clerk session cookie found")
# For custom domains, we'll also trust that Clerk redirected here
if not user_authenticated:
logger.info("Trusting Clerk redirect for custom domain flow")
user_authenticated = True
auth_method = "trusted_redirect"
oauth_session["auth_method"] = "trusted_redirect"
logger.info(f"User authenticated: {user_authenticated}, method: {auth_method}")
# Generate simple authorization code for custom domain flow
auth_code = f"clerk_custom_{session_id}_{int(time.time())}"
# Store the code mapping for token exchange
code_data = {
"session_id": session_id,
"clerk_authenticated": user_authenticated,
"auth_method": auth_method,
"custom_domain_flow": True,
"created_at": time.time(),
"expires_at": (datetime.utcnow() + timedelta(minutes=5)).timestamp(),
}
if "user_id" in oauth_session:
code_data["user_id"] = oauth_session["user_id"]
oauth_provider.storage.set_session(f"code_{auth_code}", code_data)
# Build redirect URL back to Claude
redirect_params = {
"code": auth_code,
"state": original_state
}
redirect_url = f"{oauth_session['redirect_uri']}?{urlencode(redirect_params)}"
logger.info(f"Redirecting back to Claude: {redirect_url}")
return RedirectResponse(url=redirect_url)
except Exception as e:
logger.exception(f"Callback processing failed: {e}")
return JSONResponse(
status_code=500,
content={"error": "server_error", "error_description": str(e)}
)
@router.post("/register")
async def register_client(request: Request):
"""Dynamic Client Registration (RFC 7591)"""
data = await request.json()
logger.info(f"Client registration request: {data}")
# Simple dynamic registration - accept any client
client_id = f"mcp-client-{os.urandom(8).hex()}"
return JSONResponse({
"client_id": client_id,
"client_secret": None, # Public client
"redirect_uris": data.get("redirect_uris", []),
"grant_types": ["authorization_code", "refresh_token"],
"response_types": ["code"],
"client_name": data.get("client_name", "MCP Client"),
"token_endpoint_auth_method": "none",
"client_id_issued_at": int(datetime.now().timestamp())
})
@router.post("/token")
async def token_endpoint(request: Request):
"""OAuth 2.1 Token Endpoint"""
# Parse form data
form_data = await request.form()
grant_type = form_data.get("grant_type")
code = form_data.get("code")
redirect_uri = form_data.get("redirect_uri")
client_id = form_data.get("client_id")
code_verifier = form_data.get("code_verifier")
logger.info(f"Token exchange - grant_type: {grant_type}, code: {code[:20] if code else 'None'}...")
if grant_type != "authorization_code":
return JSONResponse(
status_code=400,
content={"error": "unsupported_grant_type"}
)
try:
# OAuth token exchange - validate code and return Clerk JWT
# This supports proper OAuth flow while using Clerk JWT tokens
if not code or not redirect_uri:
logger.error("Missing required parameters: code or redirect_uri")
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "Missing code or redirect_uri"}
)
# Validate OAuth code with Clerk
if CLERK_AVAILABLE:
try:
clerk = Clerk(bearer_auth=os.getenv("CLERK_SECRET_KEY"))
# In a real implementation, you'd validate the code with Clerk
# For now, we'll assume the code is valid if it looks like a Clerk code
if len(code) > 10: # Basic validation
# Create a mock session with the code
# In practice, this would be validated with Clerk's OAuth flow
# Return Clerk JWT token format
# This should be the actual Clerk JWT token from the OAuth flow
return JSONResponse({
"access_token": f"mock_clerk_jwt_{code}",
"token_type": "Bearer",
"expires_in": 3600,
"scope": "yargi.read yargi.search"
})
else:
logger.error(f"Invalid code format: {code}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Invalid authorization code"}
)
except Exception as e:
logger.error(f"Clerk validation failed: {e}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Authorization code validation failed"}
)
else:
logger.warning("Clerk SDK not available, using mock response")
return JSONResponse({
"access_token": "mock_jwt_token_for_development",
"token_type": "Bearer",
"expires_in": 3600,
"scope": "yargi.read yargi.search"
})
except Exception as e:
logger.exception(f"Token exchange failed: {e}")
return JSONResponse(
status_code=500,
content={"error": "server_error", "error_description": str(e)}
)
@@ -1,522 +0,0 @@
"""
Simplified MCP OAuth HTTP adapter - only Clerk JWT based authentication
Uses Redis for authorization code storage to support multi-machine deployment
"""
import os
import logging
from typing import Optional
from urllib.parse import urlencode, quote
from fastapi import APIRouter, Request, Query, HTTPException
from fastapi.responses import RedirectResponse, JSONResponse
# Import Redis session store
from redis_session_store import get_redis_store
# Try to import Clerk SDK
try:
from clerk_backend_api import Clerk
CLERK_AVAILABLE = True
except ImportError:
CLERK_AVAILABLE = False
Clerk = None
logger = logging.getLogger(__name__)
router = APIRouter()
# OAuth configuration
BASE_URL = os.getenv("BASE_URL", "https://api.yargimcp.com")
CLERK_DOMAIN = os.getenv("CLERK_DOMAIN", "accounts.yargimcp.com")
# Initialize Redis store
redis_store = None
def get_redis_session_store():
"""Get Redis store instance with lazy initialization."""
global redis_store
if redis_store is None:
try:
import concurrent.futures
import functools
# Use thread pool with timeout to prevent hanging
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as executor:
future = executor.submit(get_redis_store)
try:
# 5 second timeout for Redis initialization
redis_store = future.result(timeout=5.0)
if redis_store:
logger.info("Redis session store initialized for OAuth handler")
else:
logger.warning("Redis store initialization returned None")
except concurrent.futures.TimeoutError:
logger.error("Redis initialization timed out after 5 seconds")
redis_store = None
future.cancel() # Try to cancel the hanging operation
except Exception as e:
logger.error(f"Failed to initialize Redis store: {e}")
redis_store = None
if redis_store is None:
# Fall back to in-memory storage with warning
logger.warning("Falling back to in-memory storage - multi-machine deployment will not work")
return redis_store
@router.get("/.well-known/oauth-authorization-server")
async def get_oauth_metadata():
"""OAuth 2.0 Authorization Server Metadata (RFC 8414)"""
return JSONResponse({
"issuer": BASE_URL,
"authorization_endpoint": "https://yargimcp.com/mcp-callback",
"token_endpoint": f"{BASE_URL}/token",
"registration_endpoint": f"{BASE_URL}/register",
"response_types_supported": ["code"],
"grant_types_supported": ["authorization_code"],
"code_challenge_methods_supported": ["S256"],
"token_endpoint_auth_methods_supported": ["none"],
"scopes_supported": ["read", "search", "openid", "profile", "email"],
"service_documentation": f"{BASE_URL}/mcp/"
})
@router.get("/auth/login")
async def oauth_authorize(
request: Request,
client_id: str = Query(...),
redirect_uri: str = Query(...),
response_type: str = Query("code"),
scope: Optional[str] = Query("read search"),
state: Optional[str] = Query(None),
code_challenge: Optional[str] = Query(None),
code_challenge_method: Optional[str] = Query(None)
):
"""OAuth 2.1 Authorization Endpoint - redirects to Clerk"""
logger.info(f"OAuth authorize request - client_id: {client_id}")
logger.info(f"Redirect URI: {redirect_uri}")
logger.info(f"State: {state}")
logger.info(f"PKCE Challenge: {bool(code_challenge)}")
try:
# Build callback URL with all necessary parameters
callback_url = f"{BASE_URL}/auth/callback"
callback_params = {
"client_id": client_id,
"redirect_uri": redirect_uri,
"state": state or "",
"scope": scope or "read search"
}
# Add PKCE parameters if present
if code_challenge:
callback_params["code_challenge"] = code_challenge
callback_params["code_challenge_method"] = code_challenge_method or "S256"
# Encode callback URL as redirect_url for Clerk
callback_with_params = f"{callback_url}?{urlencode(callback_params)}"
# Build Clerk sign-in URL - use yargimcp.com frontend for JWT token generation
clerk_params = {
"redirect_url": callback_with_params
}
# Use frontend sign-in page that handles JWT token generation
clerk_signin_url = f"https://yargimcp.com/sign-in?{urlencode(clerk_params)}"
logger.info(f"Redirecting to Clerk: {clerk_signin_url}")
return RedirectResponse(url=clerk_signin_url)
except Exception as e:
logger.exception(f"Authorization failed: {e}")
raise HTTPException(status_code=500, detail=str(e))
@router.get("/auth/callback")
async def oauth_callback(
request: Request,
client_id: str = Query(...),
redirect_uri: str = Query(...),
state: Optional[str] = Query(None),
scope: Optional[str] = Query("read search"),
code_challenge: Optional[str] = Query(None),
code_challenge_method: Optional[str] = Query(None),
clerk_token: Optional[str] = Query(None)
):
"""OAuth callback from Clerk - generates authorization code"""
logger.info(f"OAuth callback - client_id: {client_id}")
logger.info(f"Clerk token provided: {bool(clerk_token)}")
try:
# Validate user with Clerk and generate real JWT token
user_authenticated = False
user_id = None
session_id = None
real_jwt_token = None
if clerk_token and CLERK_AVAILABLE:
try:
# Extract user info from JWT token (no Clerk session verification needed)
import jwt
decoded_token = jwt.decode(clerk_token, options={"verify_signature": False})
user_id = decoded_token.get("user_id") or decoded_token.get("sub")
user_email = decoded_token.get("email")
token_scopes = decoded_token.get("scopes", ["read", "search"])
logger.info(f"JWT token claims - user_id: {user_id}, email: {user_email}, scopes: {token_scopes}")
if user_id and user_email:
# JWT token is already signed by Clerk and contains valid user info
user_authenticated = True
logger.info(f"User authenticated via JWT token - user_id: {user_id}")
# Use the JWT token directly as the real token (it's already from Clerk template)
real_jwt_token = clerk_token
logger.info("Using Clerk JWT token directly (already real token)")
else:
logger.error(f"Missing required fields in JWT token - user_id: {bool(user_id)}, email: {bool(user_email)}")
except Exception as e:
logger.error(f"JWT validation failed: {e}")
# Fallback to cookie validation
if not user_authenticated:
clerk_session = request.cookies.get("__session")
if clerk_session:
user_authenticated = True
logger.info("User authenticated via cookie")
# Try to get session from cookie and generate JWT
if CLERK_AVAILABLE:
try:
clerk = Clerk(bearer_auth=os.getenv("CLERK_SECRET_KEY"))
# Note: sessions.verify_session is deprecated, but we'll try
# In practice, you'd need to extract session_id from cookie
logger.info("Cookie authentication - JWT generation not implemented yet")
except Exception as e:
logger.warning(f"Failed to generate JWT from cookie: {e}")
# Only generate authorization code if we have a real JWT token
if user_authenticated and real_jwt_token:
# Generate authorization code
auth_code = f"clerk_auth_{os.urandom(16).hex()}"
# Prepare code data
import time
code_data = {
"user_id": user_id,
"session_id": session_id,
"real_jwt_token": real_jwt_token,
"user_authenticated": user_authenticated,
"client_id": client_id,
"redirect_uri": redirect_uri,
"scope": scope or "read search"
}
# Try to store in Redis, fall back to in-memory if Redis unavailable
store = get_redis_session_store()
if store:
# Store in Redis with automatic expiration
success = store.set_oauth_code(auth_code, code_data)
if success:
logger.info(f"Stored authorization code {auth_code[:10]}... in Redis with real JWT token")
else:
logger.error(f"Failed to store authorization code in Redis, falling back to in-memory")
# Fall back to in-memory storage
if not hasattr(oauth_callback, '_code_storage'):
oauth_callback._code_storage = {}
oauth_callback._code_storage[auth_code] = code_data
else:
# Fall back to in-memory storage
logger.warning("Redis not available, using in-memory storage")
if not hasattr(oauth_callback, '_code_storage'):
oauth_callback._code_storage = {}
oauth_callback._code_storage[auth_code] = code_data
logger.info(f"Stored authorization code in memory (fallback)")
# Redirect back to client with authorization code
redirect_params = {
"code": auth_code,
"state": state or ""
}
final_redirect_url = f"{redirect_uri}?{urlencode(redirect_params)}"
logger.info(f"Redirecting back to client: {final_redirect_url}")
return RedirectResponse(url=final_redirect_url)
else:
# No JWT token yet - redirect back to sign-in page to wait for authentication
logger.info("No JWT token provided - redirecting back to sign-in to complete authentication")
# Keep the same redirect URL so the flow continues
sign_in_params = {
"redirect_url": f"{request.url._url}" # Current callback URL with all params
}
sign_in_url = f"https://yargimcp.com/sign-in?{urlencode(sign_in_params)}"
logger.info(f"Redirecting back to sign-in: {sign_in_url}")
return RedirectResponse(url=sign_in_url)
except Exception as e:
logger.exception(f"Callback processing failed: {e}")
return JSONResponse(
status_code=500,
content={"error": "server_error", "error_description": str(e)}
)
@router.post("/auth/register")
async def register_client(request: Request):
"""Dynamic Client Registration (RFC 7591)"""
data = await request.json()
logger.info(f"Client registration request: {data}")
# Simple dynamic registration - accept any client
client_id = f"mcp-client-{os.urandom(8).hex()}"
return JSONResponse({
"client_id": client_id,
"client_secret": None, # Public client
"redirect_uris": data.get("redirect_uris", []),
"grant_types": ["authorization_code"],
"response_types": ["code"],
"client_name": data.get("client_name", "MCP Client"),
"token_endpoint_auth_method": "none"
})
@router.post("/auth/callback")
async def oauth_callback_post(request: Request):
"""OAuth callback POST endpoint for token exchange"""
# Parse form data (standard OAuth token exchange format)
form_data = await request.form()
grant_type = form_data.get("grant_type")
code = form_data.get("code")
redirect_uri = form_data.get("redirect_uri")
client_id = form_data.get("client_id")
code_verifier = form_data.get("code_verifier")
logger.info(f"OAuth callback POST - grant_type: {grant_type}")
logger.info(f"Code: {code[:20] if code else 'None'}...")
logger.info(f"Client ID: {client_id}")
logger.info(f"PKCE verifier: {bool(code_verifier)}")
if grant_type != "authorization_code":
return JSONResponse(
status_code=400,
content={"error": "unsupported_grant_type"}
)
if not code or not redirect_uri:
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "Missing code or redirect_uri"}
)
try:
# Validate authorization code
if not code.startswith("clerk_auth_"):
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Invalid authorization code"}
)
# Retrieve stored JWT token using authorization code from Redis or in-memory fallback
stored_code_data = None
# Try to get from Redis first, then fall back to in-memory
store = get_redis_session_store()
if store:
stored_code_data = store.get_oauth_code(code, delete_after_use=True)
if stored_code_data:
logger.info(f"Retrieved authorization code {code[:10]}... from Redis")
else:
logger.warning(f"Authorization code {code[:10]}... not found in Redis")
# Fall back to in-memory storage if Redis unavailable or code not found
if not stored_code_data and hasattr(oauth_callback, '_code_storage'):
stored_code_data = oauth_callback._code_storage.get(code)
if stored_code_data:
# Clean up in-memory storage
oauth_callback._code_storage.pop(code, None)
logger.info(f"Retrieved authorization code {code[:10]}... from in-memory storage")
if not stored_code_data:
logger.error(f"No stored data found for authorization code: {code}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Authorization code not found or expired"}
)
# Note: Redis TTL handles expiration automatically, but check for manual expiration for in-memory fallback
import time
expires_at = stored_code_data.get("expires_at", 0)
if expires_at and time.time() > expires_at:
logger.error(f"Authorization code expired: {code}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Authorization code expired"}
)
# Get the real JWT token
real_jwt_token = stored_code_data.get("real_jwt_token")
if real_jwt_token:
logger.info("Returning real Clerk JWT token")
# Note: Code already deleted from Redis, clean up in-memory fallback if used
if hasattr(oauth_callback, '_code_storage'):
oauth_callback._code_storage.pop(code, None)
return JSONResponse({
"access_token": real_jwt_token,
"token_type": "Bearer",
"expires_in": 3600,
"scope": "read search"
})
else:
logger.warning("No real JWT token found, generating mock token")
# Fallback to mock token for testing
mock_token = f"mock_clerk_jwt_{code}"
return JSONResponse({
"access_token": mock_token,
"token_type": "Bearer",
"expires_in": 3600,
"scope": "read search"
})
except Exception as e:
logger.exception(f"OAuth callback POST failed: {e}")
return JSONResponse(
status_code=500,
content={"error": "server_error", "error_description": str(e)}
)
@router.post("/register")
async def register_client(request: Request):
"""Dynamic Client Registration (RFC 7591)"""
data = await request.json()
logger.info(f"Client registration request: {data}")
# Simple dynamic registration - accept any client
client_id = f"mcp-client-{os.urandom(8).hex()}"
return JSONResponse({
"client_id": client_id,
"client_secret": None, # Public client
"redirect_uris": data.get("redirect_uris", []),
"grant_types": ["authorization_code"],
"response_types": ["code"],
"client_name": data.get("client_name", "MCP Client"),
"token_endpoint_auth_method": "none"
})
@router.post("/token")
async def token_endpoint(request: Request):
"""OAuth 2.1 Token Endpoint - exchanges code for Clerk JWT"""
# Parse form data
form_data = await request.form()
grant_type = form_data.get("grant_type")
code = form_data.get("code")
redirect_uri = form_data.get("redirect_uri")
client_id = form_data.get("client_id")
code_verifier = form_data.get("code_verifier")
logger.info(f"Token exchange - grant_type: {grant_type}")
logger.info(f"Code: {code[:20] if code else 'None'}...")
if grant_type != "authorization_code":
return JSONResponse(
status_code=400,
content={"error": "unsupported_grant_type"}
)
if not code or not redirect_uri:
return JSONResponse(
status_code=400,
content={"error": "invalid_request", "error_description": "Missing code or redirect_uri"}
)
try:
# Validate authorization code
if not code.startswith("clerk_auth_"):
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Invalid authorization code"}
)
# Retrieve stored JWT token using authorization code from Redis or in-memory fallback
stored_code_data = None
# Try to get from Redis first, then fall back to in-memory
store = get_redis_session_store()
if store:
stored_code_data = store.get_oauth_code(code, delete_after_use=True)
if stored_code_data:
logger.info(f"Retrieved authorization code {code[:10]}... from Redis (/token endpoint)")
else:
logger.warning(f"Authorization code {code[:10]}... not found in Redis (/token endpoint)")
# Fall back to in-memory storage if Redis unavailable or code not found
if not stored_code_data and hasattr(oauth_callback, '_code_storage'):
stored_code_data = oauth_callback._code_storage.get(code)
if stored_code_data:
# Clean up in-memory storage
oauth_callback._code_storage.pop(code, None)
logger.info(f"Retrieved authorization code {code[:10]}... from in-memory storage (/token endpoint)")
if not stored_code_data:
logger.error(f"No stored data found for authorization code: {code}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Authorization code not found or expired"}
)
# Note: Redis TTL handles expiration automatically, but check for manual expiration for in-memory fallback
import time
expires_at = stored_code_data.get("expires_at", 0)
if expires_at and time.time() > expires_at:
logger.error(f"Authorization code expired: {code}")
return JSONResponse(
status_code=400,
content={"error": "invalid_grant", "error_description": "Authorization code expired"}
)
# Get the real JWT token
real_jwt_token = stored_code_data.get("real_jwt_token")
if real_jwt_token:
logger.info("Returning real Clerk JWT token from /token endpoint")
# Note: Code already deleted from Redis, clean up in-memory fallback if used
if hasattr(oauth_callback, '_code_storage'):
oauth_callback._code_storage.pop(code, None)
return JSONResponse({
"access_token": real_jwt_token,
"token_type": "Bearer",
"expires_in": 3600,
"scope": "read search"
})
else:
logger.warning("No real JWT token found in /token endpoint, generating mock token")
# Fallback to mock token for testing
mock_token = f"mock_clerk_jwt_{code}"
return JSONResponse({
"access_token": mock_token,
"token_type": "Bearer",
"expires_in": 3600,
"scope": "read search"
})
except Exception as e:
logger.exception(f"Token exchange failed: {e}")
return JSONResponse(
status_code=500,
content={"error": "server_error", "error_description": str(e)}
)
File diff suppressed because it is too large Load Diff
-94
View File
@@ -1,94 +0,0 @@
events {
worker_connections 1024;
}
http {
upstream yargi_mcp {
server yargi-mcp:8000;
}
# Rate limiting
limit_req_zone $binary_remote_addr zone=api_limit:10m rate=10r/s;
limit_req_zone $binary_remote_addr zone=mcp_limit:10m rate=100r/s;
server {
listen 80;
server_name localhost;
# Redirect HTTP to HTTPS in production
# return 301 https://$server_name$request_uri;
# Security headers
add_header X-Content-Type-Options nosniff;
add_header X-Frame-Options DENY;
add_header X-XSS-Protection "1; mode=block";
add_header Referrer-Policy "strict-origin-when-cross-origin";
# API endpoints
location /api/ {
limit_req zone=api_limit burst=20 nodelay;
proxy_pass http://yargi_mcp;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# Timeouts
proxy_connect_timeout 60s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}
# MCP endpoint (higher rate limit)
location /mcp-server/mcp/ {
limit_req zone=mcp_limit burst=50 nodelay;
proxy_pass http://yargi_mcp;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
# WebSocket support
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
# Longer timeouts for MCP operations
proxy_connect_timeout 300s;
proxy_send_timeout 300s;
proxy_read_timeout 300s;
}
# Health check (no rate limit)
location /health {
proxy_pass http://yargi_mcp;
proxy_set_header Host $host;
}
# Root and other paths
location / {
limit_req zone=api_limit burst=10 nodelay;
proxy_pass http://yargi_mcp;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
# SSL configuration (uncomment for production)
# server {
# listen 443 ssl http2;
# server_name your-domain.com;
#
# ssl_certificate /etc/nginx/ssl/cert.pem;
# ssl_certificate_key /etc/nginx/ssl/key.pem;
# ssl_protocols TLSv1.2 TLSv1.3;
# ssl_ciphers HIGH:!aNULL:!MD5;
#
# # Include all location blocks from above
# }
}
Binary file not shown.

Before

Width:  |  Height:  |  Size: 533 KiB

@@ -1,66 +0,0 @@
[project]
name = "yargi-mcp"
version = "0.1.6"
description = "MCP Server For Turkish Legal Databases"
readme = "README.md"
requires-python = ">=3.11"
license = {text = "MIT"}
authors = [{name = "Said Surucu", email = "saidsrc@gmail.com"}]
keywords = ["mcp", "turkish-law", "legal", "yargitay", "danistay", "bddk", "kvkk", "turkish", "law", "court", "decisions"]
classifiers = [
"Development Status :: 4 - Beta",
"Intended Audience :: Legal Industry",
"Intended Audience :: Developers",
"License :: OSI Approved :: MIT License",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Topic :: Software Development :: Libraries :: Python Modules",
"Topic :: Text Processing :: Markup :: Markdown",
"Operating System :: OS Independent",
]
urls = {Homepage = "https://github.com/saidsurucu/yargi-mcp", Issues = "https://github.com/saidsurucu/yargi-mcp/issues"}
dependencies = [
"beautifulsoup4>=4.13.4",
"httpx>=0.28.1",
"markitdown[pdf]>=0.1.1",
"pydantic>=2.11.4",
"aiohttp>=3.11.18",
"playwright>=1.52.0",
"fastmcp>=2.10.5",
"pypdf>=5.5.0",
"fastapi>=0.115.14",
"PyJWT>=2.8.0",
"tiktoken>=0.5.0",
]
[project.optional-dependencies]
asgi = [
"uvicorn[standard]>=0.30.0",
"starlette>=0.37.0",
]
api = [
"fastapi>=0.115.0",
"uvicorn[standard]>=0.30.0",
]
production = [
"gunicorn>=22.0.0",
"uvicorn[standard]>=0.30.0",
]
saas = [
"clerk-backend-api>=3.0.0",
"stripe>=9.1.0",
"upstash-redis>=1.1.0",
]
[project.scripts]
yargi-mcp = "mcp_server_main:main"
[tool.setuptools]
py-modules = ["mcp_server_main", "mcp_auth_factory", "mcp_auth_http_adapter", "asgi_app", "fastapi_app", "starlette_app", "run_asgi", "stripe_webhook"]
[tool.setuptools.packages.find]
include = ["*_mcp_module", "mcp_auth"]
[build-system]
requires = ["setuptools>=65.0", "wheel"]
build-backend = "setuptools.build_meta"
-18
View File
@@ -1,18 +0,0 @@
{
"$schema": "https://railway.app/railway.schema.json",
"build": {
"builder": "NIXPACKS",
"buildCommand": "pip install -e .[asgi]"
},
"deploy": {
"startCommand": "uvicorn asgi_app:app --host 0.0.0.0 --port $PORT",
"healthcheckPath": "/health",
"healthcheckTimeout": 30,
"restartPolicyType": "ON_FAILURE",
"restartPolicyMaxRetries": 3
},
"variables": {
"ALLOWED_ORIGINS": "*",
"LOG_LEVEL": "info"
}
}
@@ -1,464 +0,0 @@
"""
Redis Session Store for OAuth Authorization Codes and User Sessions
This module provides Redis-based storage for OAuth authorization codes and user sessions,
enabling multi-machine deployment support by replacing in-memory storage.
Uses Upstash Redis via REST API for serverless-friendly operation.
"""
import os
import json
import time
import logging
from typing import Optional, Dict, Any, Union
from datetime import datetime, timedelta
logger = logging.getLogger(__name__)
try:
from upstash_redis import Redis
UPSTASH_AVAILABLE = True
except ImportError:
UPSTASH_AVAILABLE = False
Redis = None
# Use standard Python exceptions for Redis connection errors
import socket
from requests.exceptions import ConnectionError as RequestsConnectionError, Timeout as RequestsTimeout
class RedisSessionStore:
"""
Redis-based session store for OAuth flows and user sessions.
Uses Upstash Redis REST API for connection-free operation suitable for
multi-instance deployments on platforms like Fly.io.
"""
def __init__(self):
"""Initialize Redis connection using environment variables."""
if not UPSTASH_AVAILABLE:
raise ImportError("upstash-redis package is required. Install with: pip install upstash-redis")
# Initialize Upstash Redis client from environment with optimized connection settings
try:
# Get Upstash Redis configuration
redis_url = os.getenv("UPSTASH_REDIS_REST_URL")
redis_token = os.getenv("UPSTASH_REDIS_REST_TOKEN")
if not redis_url or not redis_token:
raise ValueError("UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN must be set")
logger.info(f"Connecting to Upstash Redis at {redis_url[:30]}...")
# Initialize with explicit configuration for better SSL handling
self.redis = Redis(
url=redis_url,
token=redis_token
)
logger.info("Upstash Redis client created")
# Skip connection test during initialization to prevent server hang
# Connection will be tested during first actual operation
logger.info("Redis client initialized - connection will be tested on first use")
except Exception as e:
logger.error(f"Failed to initialize Upstash Redis: {e}")
raise
# TTL values (in seconds)
self.oauth_code_ttl = int(os.getenv("OAUTH_CODE_TTL", "600")) # 10 minutes
self.session_ttl = int(os.getenv("SESSION_TTL", "3600")) # 1 hour
def _serialize_data(self, data: Dict[str, Any]) -> Dict[str, str]:
"""Convert data to Redis-compatible string format."""
serialized = {}
for key, value in data.items():
if isinstance(value, (dict, list)):
serialized[key] = json.dumps(value)
elif isinstance(value, (int, float)):
serialized[key] = str(value)
elif isinstance(value, bool):
serialized[key] = "true" if value else "false"
else:
serialized[key] = str(value)
return serialized
def _deserialize_data(self, data: Dict[str, str]) -> Dict[str, Any]:
"""Convert Redis string data back to original types."""
if not data:
return {}
deserialized = {}
for key, value in data.items():
if not isinstance(value, str):
deserialized[key] = value
continue
# Try to deserialize JSON
if value.startswith(('[', '{')):
try:
deserialized[key] = json.loads(value)
continue
except json.JSONDecodeError:
pass
# Try to convert numbers
if value.isdigit():
deserialized[key] = int(value)
continue
if value.replace('.', '').isdigit():
try:
deserialized[key] = float(value)
continue
except ValueError:
pass
# Handle booleans
if value in ("true", "false"):
deserialized[key] = value == "true"
continue
# Keep as string
deserialized[key] = value
return deserialized
# OAuth Authorization Code Methods
def set_oauth_code(self, code: str, data: Dict[str, Any]) -> bool:
"""
Store OAuth authorization code with automatic expiration.
Args:
code: Authorization code string
data: Code data including user_id, client_id, etc.
Returns:
True if stored successfully, False otherwise
"""
try:
key = f"oauth:code:{code}"
# Add timestamp for debugging
data_with_timestamp = data.copy()
data_with_timestamp.update({
"created_at": time.time(),
"expires_at": time.time() + self.oauth_code_ttl
})
# Serialize and store - Upstash Redis doesn't support mapping parameter
serialized_data = self._serialize_data(data_with_timestamp)
# Use individual hset calls for each field with retry logic
max_retries = 3
for attempt in range(max_retries):
try:
# Clear any existing data first
self.redis.delete(key)
# Set all fields in a pipeline-like manner
for field, value in serialized_data.items():
self.redis.hset(key, field, value)
# Set expiration
self.redis.expire(key, self.oauth_code_ttl)
logger.info(f"Stored OAuth code {code[:10]}... with TTL {self.oauth_code_ttl}s (attempt {attempt + 1})")
return True
except (RequestsConnectionError, RequestsTimeout, OSError, socket.error) as e:
logger.warning(f"Redis connection error on attempt {attempt + 1}: {e}")
if attempt == max_retries - 1:
raise # Re-raise on final attempt
time.sleep(0.5 * (attempt + 1)) # Exponential backoff
except Exception as e:
logger.error(f"Failed to store OAuth code {code[:10]}... after {max_retries} attempts: {e}")
return False
def get_oauth_code(self, code: str, delete_after_use: bool = True) -> Optional[Dict[str, Any]]:
"""
Retrieve OAuth authorization code data.
Args:
code: Authorization code string
delete_after_use: If True, delete the code after retrieval (one-time use)
Returns:
Code data dictionary or None if not found/expired
"""
max_retries = 3
for attempt in range(max_retries):
try:
key = f"oauth:code:{code}"
# Get all hash fields with retry
data = self.redis.hgetall(key)
if not data:
logger.warning(f"OAuth code {code[:10]}... not found or expired (attempt {attempt + 1})")
return None
# Deserialize data
deserialized_data = self._deserialize_data(data)
# Check manual expiration (in case Redis TTL failed)
expires_at = deserialized_data.get("expires_at", 0)
if expires_at and time.time() > expires_at:
logger.warning(f"OAuth code {code[:10]}... manually expired")
try:
self.redis.delete(key)
except Exception as del_error:
logger.warning(f"Failed to delete expired code: {del_error}")
return None
# Delete after use for security (one-time use)
if delete_after_use:
try:
self.redis.delete(key)
logger.info(f"Retrieved and deleted OAuth code {code[:10]}... (attempt {attempt + 1})")
except Exception as del_error:
logger.warning(f"Failed to delete code after use: {del_error}")
# Continue anyway since we got the data
else:
logger.info(f"Retrieved OAuth code {code[:10]}... (not deleted, attempt {attempt + 1})")
return deserialized_data
except (RequestsConnectionError, RequestsTimeout, OSError, socket.error) as e:
logger.warning(f"Redis connection error on retrieval attempt {attempt + 1}: {e}")
if attempt == max_retries - 1:
logger.error(f"Failed to retrieve OAuth code {code[:10]}... after {max_retries} attempts: {e}")
return None
time.sleep(0.5 * (attempt + 1)) # Exponential backoff
except Exception as e:
logger.error(f"Failed to retrieve OAuth code {code[:10]}... on attempt {attempt + 1}: {e}")
if attempt == max_retries - 1:
return None
time.sleep(0.5 * (attempt + 1))
return None
# User Session Methods
def set_session(self, session_id: str, user_data: Dict[str, Any]) -> bool:
"""
Store user session data with sliding expiration.
Args:
session_id: Unique session identifier
user_data: User session data (user_id, email, scopes, etc.)
Returns:
True if stored successfully, False otherwise
"""
try:
key = f"session:{session_id}"
# Add session metadata
session_data = user_data.copy()
session_data.update({
"session_id": session_id,
"created_at": time.time(),
"last_accessed": time.time()
})
# Serialize and store - Upstash Redis doesn't support mapping parameter
serialized_data = self._serialize_data(session_data)
# Use individual hset calls for each field (Upstash compatibility)
for field, value in serialized_data.items():
self.redis.hset(key, field, value)
self.redis.expire(key, self.session_ttl)
logger.info(f"Stored session {session_id[:10]}... with TTL {self.session_ttl}s")
return True
except Exception as e:
logger.error(f"Failed to store session {session_id[:10]}...: {e}")
return False
def get_session(self, session_id: str, refresh_ttl: bool = True) -> Optional[Dict[str, Any]]:
"""
Retrieve user session data.
Args:
session_id: Session identifier
refresh_ttl: If True, extend session TTL on access
Returns:
Session data dictionary or None if not found/expired
"""
try:
key = f"session:{session_id}"
# Get session data
data = self.redis.hgetall(key)
if not data:
logger.warning(f"Session {session_id[:10]}... not found or expired")
return None
# Deserialize data
session_data = self._deserialize_data(data)
# Update last accessed time and refresh TTL
if refresh_ttl:
session_data["last_accessed"] = time.time()
self.redis.hset(key, "last_accessed", str(time.time()))
self.redis.expire(key, self.session_ttl)
logger.debug(f"Refreshed session {session_id[:10]}... TTL")
return session_data
except Exception as e:
logger.error(f"Failed to retrieve session {session_id[:10]}...: {e}")
return None
def delete_session(self, session_id: str) -> bool:
"""
Delete user session (logout).
Args:
session_id: Session identifier
Returns:
True if deleted successfully, False otherwise
"""
try:
key = f"session:{session_id}"
result = self.redis.delete(key)
if result:
logger.info(f"Deleted session {session_id[:10]}...")
return True
else:
logger.warning(f"Session {session_id[:10]}... not found for deletion")
return False
except Exception as e:
logger.error(f"Failed to delete session {session_id[:10]}...: {e}")
return False
# Health Check Methods
def health_check(self) -> Dict[str, Any]:
"""
Perform Redis health check.
Returns:
Health status dictionary
"""
try:
# Test basic operations
test_key = f"health:check:{int(time.time())}"
test_value = {"timestamp": time.time(), "test": True}
# Test set - Use individual hset calls for Upstash compatibility
serialized_test = self._serialize_data(test_value)
for field, value in serialized_test.items():
self.redis.hset(test_key, field, value)
# Test get
retrieved = self.redis.hgetall(test_key)
# Test delete
self.redis.delete(test_key)
return {
"status": "healthy",
"redis_connected": True,
"operations_working": bool(retrieved),
"timestamp": datetime.utcnow().isoformat()
}
except Exception as e:
logger.error(f"Redis health check failed: {e}")
return {
"status": "unhealthy",
"redis_connected": False,
"error": str(e),
"timestamp": datetime.utcnow().isoformat()
}
def get_stats(self) -> Dict[str, Any]:
"""
Get Redis usage statistics.
Returns:
Statistics dictionary
"""
try:
# Get basic info (not all Upstash plans support INFO command)
stats = {
"oauth_codes_pattern": "oauth:code:*",
"sessions_pattern": "session:*",
"timestamp": datetime.utcnow().isoformat()
}
try:
# Try to get counts (may fail on some Upstash plans)
oauth_keys = self.redis.keys("oauth:code:*")
session_keys = self.redis.keys("session:*")
stats.update({
"active_oauth_codes": len(oauth_keys) if oauth_keys else 0,
"active_sessions": len(session_keys) if session_keys else 0
})
except Exception as e:
logger.warning(f"Could not get detailed stats: {e}")
stats["warning"] = "Detailed stats not available on this Redis plan"
return stats
except Exception as e:
logger.error(f"Failed to get Redis stats: {e}")
return {"error": str(e), "timestamp": datetime.utcnow().isoformat()}
# Global instance for easy importing
redis_store = None
def get_redis_store() -> Optional[RedisSessionStore]:
"""
Get global Redis store instance (singleton pattern).
Returns:
RedisSessionStore instance or None if initialization fails
"""
global redis_store
if redis_store is None:
try:
logger.info("Initializing Redis store...")
redis_store = RedisSessionStore()
logger.info("Redis store initialized successfully")
except Exception as e:
logger.error(f"Failed to initialize Redis store: {e}")
redis_store = None
return redis_store
def init_redis_store() -> RedisSessionStore:
"""
Initialize Redis store and perform health check.
Returns:
RedisSessionStore instance
Raises:
Exception if Redis is not available or unhealthy
"""
store = get_redis_store()
# Perform health check
health = store.health_check()
if health["status"] != "healthy":
raise Exception(f"Redis health check failed: {health}")
logger.info("Redis session store initialized and healthy")
return store
@@ -1,407 +0,0 @@
# rekabet_mcp_module/client.py
import httpx
from bs4 import BeautifulSoup
from typing import List, Optional, Tuple, Dict, Any
import logging
import html
import re
import io # For io.BytesIO
from urllib.parse import urlencode, urljoin, quote, parse_qs, urlparse
from markitdown import MarkItDown
import math
# pypdf for PDF processing (lighter alternative to PyMuPDF)
from pypdf import PdfReader, PdfWriter # PyPDF2'nin devamı niteliğindeki pypdf
from .models import (
RekabetKurumuSearchRequest,
RekabetDecisionSummary,
RekabetSearchResult,
RekabetDocument,
RekabetKararTuruGuidEnum
)
from pydantic import HttpUrl # Ensure HttpUrl is imported from pydantic
logger = logging.getLogger(__name__)
if not logger.hasHandlers(): # Pragma: no cover
logging.basicConfig(
level=logging.INFO, # Varsayılan log seviyesi
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
# Debug betiğinde daha detaylı loglama için seviye ayrıca ayarlanabilir.
class RekabetKurumuApiClient:
BASE_URL = "https://www.rekabet.gov.tr"
SEARCH_PATH = "/tr/Kararlar"
DECISION_LANDING_PATH_TEMPLATE = "/Karar"
# PDF sayfa bazlı Markdown döndürüldüğü için bu sabit artık doğrudan kullanılmıyor.
# DOCUMENT_MARKDOWN_CHUNK_SIZE = 5000
def __init__(self, request_timeout: float = 60.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
},
timeout=request_timeout,
verify=True,
follow_redirects=True
)
def _build_search_query_params(self, params: RekabetKurumuSearchRequest) -> List[Tuple[str, str]]:
query_params: List[Tuple[str, str]] = []
query_params.append(("sayfaAdi", params.sayfaAdi if params.sayfaAdi is not None else ""))
query_params.append(("YayinlanmaTarihi", params.YayinlanmaTarihi if params.YayinlanmaTarihi is not None else ""))
query_params.append(("PdfText", params.PdfText if params.PdfText is not None else ""))
karar_turu_id_value = ""
if params.KararTuruID is not None:
karar_turu_id_value = params.KararTuruID.value if params.KararTuruID.value != "ALL" else ""
query_params.append(("KararTuruID", karar_turu_id_value))
query_params.append(("KararSayisi", params.KararSayisi if params.KararSayisi is not None else ""))
query_params.append(("KararTarihi", params.KararTarihi if params.KararTarihi is not None else ""))
if params.page and params.page > 1:
query_params.append(("page", str(params.page)))
return query_params
async def search_decisions(self, params: RekabetKurumuSearchRequest) -> RekabetSearchResult:
request_path = self.SEARCH_PATH
final_query_params = self._build_search_query_params(params)
logger.info(f"RekabetKurumuApiClient: Performing search. Path: {request_path}, Parameters: {final_query_params}")
try:
response = await self.http_client.get(request_path, params=final_query_params)
response.raise_for_status()
html_content = response.text
except httpx.RequestError as e:
logger.error(f"RekabetKurumuApiClient: HTTP request error during search: {e}")
raise
soup = BeautifulSoup(html_content, 'html.parser')
processed_decisions: List[RekabetDecisionSummary] = []
total_records: Optional[int] = None
total_pages: Optional[int] = None
pagination_div = soup.find("div", class_="yazi01")
if pagination_div:
text_content = pagination_div.get_text(separator=" ", strip=True)
total_match = re.search(r"Toplam\s*:\s*(\d+)", text_content)
if total_match:
try:
total_records = int(total_match.group(1))
logger.debug(f"Total records found from pagination: {total_records}")
except ValueError:
logger.warning(f"Could not convert 'Toplam' value to int: {total_match.group(1)}")
else:
logger.warning("'Toplam :' string not found in pagination section.")
results_per_page_assumed = 10
if total_records is not None:
calculated_total_pages = math.ceil(total_records / results_per_page_assumed)
total_pages = calculated_total_pages if calculated_total_pages > 0 else (1 if total_records > 0 else 0)
logger.debug(f"Calculated total pages: {total_pages}")
if total_pages is None: # Fallback if total_records couldn't be parsed
last_page_link = pagination_div.select_one("li.PagedList-skipToLast a")
if last_page_link and last_page_link.has_attr('href'):
qs = parse_qs(urlparse(last_page_link['href']).query)
if 'page' in qs and qs['page']:
try:
total_pages = int(qs['page'][0])
logger.debug(f"Total pages found from 'Last >>' link: {total_pages}")
except ValueError:
logger.warning(f"Could not convert page value from 'Last >>' link to int: {qs['page'][0]}")
elif total_records == 0 : total_pages = 0 # If no records, 0 pages
elif total_records is not None and total_records > 0 : total_pages = 1 # If records exist but no last page link (e.g. single page)
else: logger.warning("'Last >>' link not found in pagination section.")
decision_tables_container = soup.find("div", id="kararList")
if not decision_tables_container:
logger.warning("`div#kararList` (decision list container) not found. HTML structure might have changed or no decisions on this page.")
else:
decision_tables = decision_tables_container.find_all("table", class_="equalDivide")
logger.info(f"Found {len(decision_tables)} 'table' elements with class='equalDivide' for parsing.")
if not decision_tables and total_records is not None and total_records > 0 :
logger.warning(f"Page indicates {total_records} records but no decision tables found with class='equalDivide'.")
for idx, table in enumerate(decision_tables):
logger.debug(f"Processing table {idx + 1}...")
try:
rows = table.find_all("tr")
if len(rows) != 3:
logger.warning(f"Table {idx + 1} has an unexpected number of rows ({len(rows)} instead of 3). Skipping. HTML snippet:\n{table.prettify()[:500]}")
continue
# Row 1: Publication Date, Decision Number, Related Cases Link
td_elements_r1 = rows[0].find_all("td")
pub_date = td_elements_r1[0].get_text(strip=True) if len(td_elements_r1) > 0 else None
dec_num = td_elements_r1[1].get_text(strip=True) if len(td_elements_r1) > 1 else None
related_cases_link_tag = td_elements_r1[2].find("a", href=True) if len(td_elements_r1) > 2 else None
related_cases_url_str: Optional[str] = None
karar_id_from_related: Optional[str] = None
if related_cases_link_tag and related_cases_link_tag.has_attr('href'):
related_cases_url_str = urljoin(self.BASE_URL, related_cases_link_tag['href'])
qs_related = parse_qs(urlparse(related_cases_link_tag['href']).query)
if 'kararId' in qs_related and qs_related['kararId']:
karar_id_from_related = qs_related['kararId'][0]
# Row 2: Decision Date, Decision Type
td_elements_r2 = rows[1].find_all("td")
dec_date = td_elements_r2[0].get_text(strip=True) if len(td_elements_r2) > 0 else None
dec_type_text = td_elements_r2[1].get_text(strip=True) if len(td_elements_r2) > 1 else None
# Row 3: Title and Main Decision Link
title_cell = rows[2].find("td", colspan="5")
decision_link_tag = title_cell.find("a", href=True) if title_cell else None
title_text: Optional[str] = None
decision_landing_url_str: Optional[str] = None
karar_id_from_main_link: Optional[str] = None
if decision_link_tag and decision_link_tag.has_attr('href'):
title_text = decision_link_tag.get_text(strip=True)
href_val = decision_link_tag['href']
if href_val.startswith(self.DECISION_LANDING_PATH_TEMPLATE + "?kararId="): # Ensure it's a decision link
decision_landing_url_str = urljoin(self.BASE_URL, href_val)
qs_main = parse_qs(urlparse(href_val).query)
if 'kararId' in qs_main and qs_main['kararId']:
karar_id_from_main_link = qs_main['kararId'][0]
else:
logger.warning(f"Table {idx+1} decision link has unexpected format: {href_val}")
else:
logger.warning(f"Table {idx+1} could not find title/decision link tag.")
current_karar_id = karar_id_from_main_link or karar_id_from_related
if not current_karar_id:
logger.warning(f"Table {idx+1} Karar ID not found. Skipping. Title (if any): {title_text}")
continue
# Convert string URLs to HttpUrl for the model
final_decision_url = HttpUrl(decision_landing_url_str) if decision_landing_url_str else None
final_related_cases_url = HttpUrl(related_cases_url_str) if related_cases_url_str else None
processed_decisions.append(RekabetDecisionSummary(
publication_date=pub_date, decision_number=dec_num, decision_date=dec_date,
decision_type_text=dec_type_text, title=title_text,
decision_url=final_decision_url,
karar_id=current_karar_id,
related_cases_url=final_related_cases_url
))
logger.debug(f"Table {idx+1} parsed successfully: Karar ID '{current_karar_id}', Title '{title_text[:50] if title_text else 'N/A'}...'")
except Exception as e:
logger.warning(f"RekabetKurumuApiClient: Error parsing decision summary {idx+1}: {e}. Problematic Table HTML:\n{table.prettify()}", exc_info=True)
continue
return RekabetSearchResult(
decisions=processed_decisions, total_records_found=total_records,
retrieved_page_number=params.page, total_pages=total_pages if total_pages is not None else 0
)
async def _extract_pdf_url_and_landing_page_metadata(self, karar_id: str, landing_page_html: str, landing_page_url: str) -> Dict[str, Any]:
soup = BeautifulSoup(landing_page_html, 'html.parser')
data: Dict[str, Any] = {
"pdf_url": None,
"title_on_landing_page": soup.title.string.strip() if soup.title and soup.title.string else f"Rekabet Kurumu Kararı {karar_id}",
}
# This part needs to be robust and specific to Rekabet Kurumu's landing page structure.
# Look for common patterns: direct links, download buttons, embedded viewers.
pdf_anchor = soup.find("a", href=re.compile(r"\.pdf(\?|$)", re.IGNORECASE)) # Basic PDF link
if not pdf_anchor: # Try other common patterns if the basic one fails
# Example: Look for links with specific text or class
pdf_anchor = soup.find("a", string=re.compile(r"karar metni|pdf indir", re.IGNORECASE))
if pdf_anchor and pdf_anchor.has_attr('href'):
pdf_path = pdf_anchor['href']
data["pdf_url"] = urljoin(landing_page_url, pdf_path)
logger.info(f"PDF link found on landing page (<a>): {data['pdf_url']}")
else:
iframe_pdf = soup.find("iframe", src=re.compile(r"\.pdf(\?|$)", re.IGNORECASE))
if iframe_pdf and iframe_pdf.has_attr('src'):
pdf_path = iframe_pdf['src']
data["pdf_url"] = urljoin(landing_page_url, pdf_path)
logger.info(f"PDF link found on landing page (<iframe>): {data['pdf_url']}")
else:
embed_pdf = soup.find("embed", src=re.compile(r"\.pdf(\?|$)", re.IGNORECASE), type="application/pdf")
if embed_pdf and embed_pdf.has_attr('src'):
pdf_path = embed_pdf['src']
data["pdf_url"] = urljoin(landing_page_url, pdf_path)
logger.info(f"PDF link found on landing page (<embed>): {data['pdf_url']}")
else:
logger.warning(f"No PDF link found on landing page {landing_page_url} for kararId {karar_id} using common selectors.")
return data
async def _download_pdf_bytes(self, pdf_url: str) -> Optional[bytes]:
try:
url_to_fetch = pdf_url if pdf_url.startswith(('http://', 'https://')) else urljoin(self.BASE_URL, pdf_url)
logger.info(f"Downloading PDF from: {url_to_fetch}")
response = await self.http_client.get(url_to_fetch)
response.raise_for_status()
pdf_bytes = await response.aread()
logger.info(f"PDF content downloaded ({len(pdf_bytes)} bytes) from: {url_to_fetch}")
return pdf_bytes
except httpx.RequestError as e:
logger.error(f"HTTP error downloading PDF from {pdf_url}: {e}")
except Exception as e:
logger.error(f"General error downloading PDF from {pdf_url}: {e}")
return None
def _extract_single_pdf_page_as_pdf_bytes(self, original_pdf_bytes: bytes, page_number_to_extract: int) -> Tuple[Optional[bytes], int]:
total_pages_in_original_pdf = 0
single_page_pdf_bytes: Optional[bytes] = None
if not original_pdf_bytes:
logger.warning("No original PDF bytes provided for page extraction.")
return None, 0
try:
pdf_stream = io.BytesIO(original_pdf_bytes)
reader = PdfReader(pdf_stream)
total_pages_in_original_pdf = len(reader.pages)
if not (0 < page_number_to_extract <= total_pages_in_original_pdf):
logger.warning(f"Requested page number ({page_number_to_extract}) is out of PDF page range (1-{total_pages_in_original_pdf}).")
return None, total_pages_in_original_pdf
writer = PdfWriter()
writer.add_page(reader.pages[page_number_to_extract - 1]) # pypdf is 0-indexed
output_pdf_stream = io.BytesIO()
writer.write(output_pdf_stream)
single_page_pdf_bytes = output_pdf_stream.getvalue()
logger.debug(f"Page {page_number_to_extract} of original PDF (total {total_pages_in_original_pdf} pages) extracted as new PDF using pypdf.")
except Exception as e:
logger.error(f"Error extracting PDF page using pypdf: {e}", exc_info=True)
return None, total_pages_in_original_pdf
return single_page_pdf_bytes, total_pages_in_original_pdf
def _convert_pdf_bytes_to_markdown(self, pdf_bytes: bytes, source_url_for_logging: str) -> Optional[str]:
if not pdf_bytes:
logger.warning(f"No PDF bytes provided for Markdown conversion (source: {source_url_for_logging}).")
return None
pdf_stream = io.BytesIO(pdf_bytes)
try:
md_converter = MarkItDown(enable_plugins=False)
conversion_result = md_converter.convert(pdf_stream)
markdown_text = conversion_result.text_content
if not markdown_text:
logger.warning(f"MarkItDown returned empty content from PDF byte stream (source: {source_url_for_logging}). PDF page might be image-based or MarkItDown could not process the PDF stream.")
return markdown_text
except Exception as e:
logger.error(f"MarkItDown conversion error for PDF byte stream (source: {source_url_for_logging}): {e}", exc_info=True)
return None
async def get_decision_document(self, karar_id: str, page_number: int = 1) -> RekabetDocument:
if not karar_id:
return RekabetDocument(
source_landing_page_url=HttpUrl(f"{self.BASE_URL}"),
karar_id=karar_id or "UNKNOWN_KARAR_ID",
error_message="karar_id is required.",
current_page=1, total_pages=0, is_paginated=False )
decision_url_path = f"{self.DECISION_LANDING_PATH_TEMPLATE}?kararId={karar_id}"
full_landing_page_url = urljoin(self.BASE_URL, decision_url_path)
logger.info(f"RekabetKurumuApiClient: Getting decision document: {full_landing_page_url}, Requested PDF Page: {page_number}")
pdf_url_to_report: Optional[HttpUrl] = None
title_to_report: Optional[str] = f"Rekabet Kurumu Kararı {karar_id}" # Default
error_message: Optional[str] = None
markdown_for_requested_page: Optional[str] = None
total_pdf_pages: int = 0
try:
async with self.http_client.stream("GET", full_landing_page_url) as response:
response.raise_for_status()
content_type = response.headers.get("content-type", "").lower()
final_url_of_response = HttpUrl(str(response.url))
original_pdf_bytes: Optional[bytes] = None
if "application/pdf" in content_type:
logger.info(f"URL {final_url_of_response} is a direct PDF. Processing content.")
pdf_url_to_report = final_url_of_response
original_pdf_bytes = await response.aread()
elif "text/html" in content_type:
logger.info(f"URL {final_url_of_response} is an HTML landing page. Looking for PDF link.")
landing_page_html_bytes = await response.aread()
detected_charset = response.charset_encoding or 'utf-8'
try: landing_page_html = landing_page_html_bytes.decode(detected_charset)
except UnicodeDecodeError: landing_page_html = landing_page_html_bytes.decode('utf-8', errors='replace')
if landing_page_html.strip():
landing_page_data = self._extract_pdf_url_and_landing_page_metadata(karar_id, landing_page_html, str(final_url_of_response))
pdf_url_str_from_html = landing_page_data.get("pdf_url")
if landing_page_data.get("title_on_landing_page"): title_to_report = landing_page_data.get("title_on_landing_page")
if pdf_url_str_from_html:
pdf_url_to_report = HttpUrl(pdf_url_str_from_html)
original_pdf_bytes = await self._download_pdf_bytes(str(pdf_url_to_report))
else: error_message = (error_message or "") + " PDF URL not found on HTML landing page."
else: error_message = "Decision landing page content is empty."
else: error_message = f"Unexpected content type ({content_type}) for URL: {final_url_of_response}"
if original_pdf_bytes:
single_page_pdf_bytes, total_pdf_pages_from_extraction = self._extract_single_pdf_page_as_pdf_bytes(original_pdf_bytes, page_number)
total_pdf_pages = total_pdf_pages_from_extraction
if single_page_pdf_bytes:
markdown_for_requested_page = self._convert_pdf_bytes_to_markdown(single_page_pdf_bytes, str(pdf_url_to_report or full_landing_page_url))
if not markdown_for_requested_page:
error_message = (error_message or "") + f"; Could not convert page {page_number} of PDF to Markdown."
elif total_pdf_pages > 0 :
error_message = (error_message or "") + f"; Could not extract page {page_number} from PDF (page may be out of range or extraction failed)."
else:
error_message = (error_message or "") + "; PDF could not be processed or page count was zero (original PDF might be invalid)."
elif not error_message:
error_message = "PDF content could not be downloaded or identified."
is_paginated = total_pdf_pages > 1
current_page_final = page_number
if total_pdf_pages > 0:
current_page_final = max(1, min(page_number, total_pdf_pages))
elif markdown_for_requested_page is None:
current_page_final = 1
# If markdown is None but there was no specific error for markdown conversion (e.g. PDF not found first)
# make sure error_message reflects that.
if markdown_for_requested_page is None and pdf_url_to_report and not error_message:
error_message = (error_message or "") + "; Failed to produce Markdown from PDF page."
return RekabetDocument(
source_landing_page_url=full_landing_page_url, karar_id=karar_id,
title_on_landing_page=title_to_report, pdf_url=pdf_url_to_report,
markdown_chunk=markdown_for_requested_page, current_page=current_page_final,
total_pages=total_pdf_pages, is_paginated=is_paginated,
error_message=error_message.strip("; ") if error_message else None )
except httpx.HTTPStatusError as e: error_msg_detail = f"HTTP Status error {e.response.status_code} while processing decision page."
except httpx.RequestError as e: error_msg_detail = f"HTTP Request error while processing decision page: {str(e)}"
except Exception as e: error_msg_detail = f"General error while processing decision: {str(e)}"
exc_info_flag = not isinstance(e, (httpx.HTTPStatusError, httpx.RequestError)) if 'e' in locals() else True
logger.error(f"RekabetKurumuApiClient: Error processing decision {karar_id} from {full_landing_page_url}: {error_msg_detail}", exc_info=exc_info_flag)
error_message = (error_message + "; " if error_message else "") + error_msg_detail
return RekabetDocument(
source_landing_page_url=full_landing_page_url, karar_id=karar_id,
title_on_landing_page=title_to_report, pdf_url=pdf_url_to_report,
markdown_chunk=None, current_page=page_number, total_pages=0, is_paginated=False,
error_message=error_message.strip("; ") if error_message else "An unexpected error occurred." )
async def close_client_session(self): # Pragma: no cover
if hasattr(self, 'http_client') and self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("RekabetKurumuApiClient: HTTP client session closed.")
@@ -1,71 +0,0 @@
# rekabet_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional, Any
from enum import Enum
# Enum for decision type GUIDs (used by the client and expected by the website)
class RekabetKararTuruGuidEnum(str, Enum):
TUMU = "ALL" # Represents "All" or "Select Decision Type"
BIRLESME_DEVRALMA = "2fff0979-9f9d-42d7-8c2e-a30705889542" # Merger and Acquisition
DIGER = "dda8feaf-c919-405c-9da1-823f22b45ad9" # Other
MENFI_TESPIT_MUAFIYET = "95ccd210-5304-49c5-b9e0-8ee53c50d4e8" # Negative Clearance and Exemption
OZELLESTIRME = "e1f14505-842b-4af5-95d1-312d6de1a541" # Privatization
REKABET_IHLALI = "720614bf-efd1-4dca-9785-b98eb65f2677" # Competition Infringement
# Enum for user-friendly decision type names (for server tool parameters)
# These correspond to the display names on the website's select dropdown.
class RekabetKararTuruAdiEnum(str, Enum):
TUMU = "Tümü" # Corresponds to the empty value "" for GUID, meaning "All"
BIRLESME_VE_DEVRALMA = "Birleşme ve Devralma"
DIGER = "Diğer"
MENFI_TESPIT_VE_MUAFIYET = "Menfi Tespit ve Muafiyet"
OZELLESTIRME = "Özelleştirme"
REKABET_IHLALI = "Rekabet İhlali"
class RekabetKurumuSearchRequest(BaseModel):
"""Model for Rekabet Kurumu (Turkish Competition Authority) search request."""
sayfaAdi: str = Field("", description="Title")
YayinlanmaTarihi: str = Field("", description="Date")
PdfText: str = Field("", description="Text")
KararTuruID: RekabetKararTuruGuidEnum = Field(RekabetKararTuruGuidEnum.TUMU, description="Type")
KararSayisi: str = Field("", description="No")
KararTarihi: str = Field("", description="Date")
page: int = Field(1, ge=1, description="Page")
class RekabetDecisionSummary(BaseModel):
"""Model for a single Rekabet Kurumu decision summary from search results."""
publication_date: str = Field("", description="Pub date")
decision_number: str = Field("", description="Number")
decision_date: str = Field("", description="Date")
decision_type_text: str = Field("", description="Type")
title: str = Field("", description="Title")
decision_url: str = Field("", description="URL")
karar_id: str = Field("", description="ID")
related_cases_url: str = Field("", description="Cases URL")
class RekabetSearchResult(BaseModel):
"""Model for the overall search result for Rekabet Kurumu decisions."""
decisions: List[RekabetDecisionSummary]
total_records_found: int = Field(0, description="Total")
retrieved_page_number: int = Field(description="Page")
total_pages: int = Field(0, description="Pages")
class RekabetDocument(BaseModel):
"""
Model for a Rekabet Kurumu decision document.
Contains metadata from the landing page, a link to the PDF,
and the PDF's content converted to paginated Markdown.
"""
source_landing_page_url: HttpUrl = Field(description="Source URL")
karar_id: str = Field(description="ID")
title_on_landing_page: Optional[str] = Field(None, description="Title")
pdf_url: Optional[HttpUrl] = Field(None, description="PDF URL")
markdown_chunk: Optional[str] = Field(None, description="Content")
current_page: int = Field(1, description="Page")
total_pages: int = Field(1, description="Total pages")
is_paginated: bool = Field(False, description="Paginated")
error_message: Optional[str] = Field(None, description="Error")
-119
View File
@@ -1,119 +0,0 @@
#!/usr/bin/env python3
"""
Standalone ASGI server runner for Yargı MCP
This script provides a simple way to run the Yargı MCP server
as a web service using uvicorn.
Usage:
python run_asgi.py
python run_asgi.py --host 0.0.0.0 --port 8080
python run_asgi.py --reload # For development
"""
import os
import sys
import argparse
import logging
from pathlib import Path
# Add project root to Python path
sys.path.insert(0, str(Path(__file__).parent))
try:
import uvicorn
except ImportError:
print("Error: uvicorn is not installed.")
print("Please install it with: pip install uvicorn")
sys.exit(1)
# Configure logging
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
def main():
parser = argparse.ArgumentParser(
description="Run Yargı MCP server as an ASGI web service"
)
parser.add_argument(
"--host",
type=str,
default=os.getenv("HOST", "127.0.0.1"),
help="Host to bind to (default: 127.0.0.1)"
)
parser.add_argument(
"--port",
type=int,
default=int(os.getenv("PORT", "8000")),
help="Port to bind to (default: 8000)"
)
parser.add_argument(
"--reload",
action="store_true",
help="Enable auto-reload for development"
)
parser.add_argument(
"--transport",
choices=["http", "sse"],
default="http",
help="Transport type (default: http)"
)
parser.add_argument(
"--log-level",
choices=["debug", "info", "warning", "error"],
default=os.getenv("LOG_LEVEL", "info").lower(),
help="Log level (default: info)"
)
parser.add_argument(
"--workers",
type=int,
default=1,
help="Number of worker processes (default: 1)"
)
args = parser.parse_args()
# Select app based on transport
app_name = "asgi_app:app" if args.transport == "http" else "asgi_app:sse_app"
# Configure uvicorn
config = {
"app": app_name,
"host": args.host,
"port": args.port,
"log_level": args.log_level,
"reload": args.reload,
"access_log": True,
}
# Add workers only if not in reload mode
if not args.reload and args.workers > 1:
config["workers"] = args.workers
# Print startup information
print(f"Starting Yargı MCP server...")
print(f"Host: {args.host}")
print(f"Port: {args.port}")
print(f"Transport: {args.transport}")
print(f"Log level: {args.log_level}")
if args.reload:
print("Auto-reload: enabled")
else:
print(f"Workers: {args.workers}")
print(f"\nServer will be available at: http://{args.host}:{args.port}")
print(f"MCP endpoint: http://{args.host}:{args.port}/mcp/")
print(f"Health check: http://{args.host}:{args.port}/health")
print(f"API status: http://{args.host}:{args.port}/status")
print("\nPress CTRL+C to stop the server\n")
# Run uvicorn
try:
uvicorn.run(**config)
except KeyboardInterrupt:
print("\nShutting down server...")
sys.exit(0)
if __name__ == "__main__":
main()
@@ -1,56 +0,0 @@
# sayistay_mcp_module/__init__.py
"""
Sayıştay (Turkish Court of Accounts) MCP Module
This module provides access to three types of Sayıştay decisions:
- Genel Kurul (General Assembly) decisions
- Temyiz Kurulu (Appeals Board) decisions
- Daire (Chamber) decisions
The module handles ASP.NET WebForms authentication with CSRF tokens
and DataTables-based pagination for comprehensive decision search.
"""
from .client import SayistayApiClient
from .models import (
# Genel Kurul models
GenelKurulSearchRequest,
GenelKurulSearchResponse,
GenelKurulDecision,
# Temyiz Kurulu models
TemyizKuruluSearchRequest,
TemyizKuruluSearchResponse,
TemyizKuruluDecision,
# Daire models
DaireSearchRequest,
DaireSearchResponse,
DaireDecision,
# Document models
SayistayDocumentMarkdown
)
from .enums import (
DaireEnum,
KamuIdaresiTuruEnum,
WebKararKonusuEnum
)
__all__ = [
"SayistayApiClient",
"GenelKurulSearchRequest",
"GenelKurulSearchResponse",
"GenelKurulDecision",
"TemyizKuruluSearchRequest",
"TemyizKuruluSearchResponse",
"TemyizKuruluDecision",
"DaireSearchRequest",
"DaireSearchResponse",
"DaireDecision",
"SayistayDocumentMarkdown",
"DaireEnum",
"KamuIdaresiTuruEnum",
"WebKararKonusuEnum"
]
@@ -1,689 +0,0 @@
# sayistay_mcp_module/client.py
import httpx
import re
from bs4 import BeautifulSoup
from typing import Dict, Any, List, Optional, Tuple
import logging
import html
import io
from urllib.parse import urlencode, urljoin
from markitdown import MarkItDown
from .models import (
GenelKurulSearchRequest, GenelKurulSearchResponse, GenelKurulDecision,
TemyizKuruluSearchRequest, TemyizKuruluSearchResponse, TemyizKuruluDecision,
DaireSearchRequest, DaireSearchResponse, DaireDecision,
SayistayDocumentMarkdown
)
from .enums import DaireEnum, KamuIdaresiTuruEnum, WebKararKonusuEnum, WEB_KARAR_KONUSU_MAPPING
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
)
class SayistayApiClient:
"""
API Client for Sayıştay (Turkish Court of Accounts) decision search system.
Handles three types of decisions:
- Genel Kurul (General Assembly): Precedent-setting interpretive decisions
- Temyiz Kurulu (Appeals Board): Appeals against chamber decisions
- Daire (Chamber): First-instance audit findings and sanctions
Features:
- ASP.NET WebForms session management with CSRF tokens
- DataTables-based pagination and filtering
- Automatic session refresh on expiration
- Document retrieval with Markdown conversion
"""
BASE_URL = "https://www.sayistay.gov.tr"
# Search endpoints for each decision type
GENEL_KURUL_ENDPOINT = "/KararlarGenelKurul/DataTablesList"
TEMYIZ_KURULU_ENDPOINT = "/KararlarTemyiz/DataTablesList"
DAIRE_ENDPOINT = "/KararlarDaire/DataTablesList"
# Page endpoints for session initialization and document access
GENEL_KURUL_PAGE = "/KararlarGenelKurul"
TEMYIZ_KURULU_PAGE = "/KararlarTemyiz"
DAIRE_PAGE = "/KararlarDaire"
def __init__(self, request_timeout: float = 60.0):
self.request_timeout = request_timeout
self.session_cookies: Dict[str, str] = {}
self.csrf_tokens: Dict[str, str] = {} # Store tokens for each endpoint
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Accept": "application/json, text/javascript, */*; q=0.01",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"Content-Type": "application/x-www-form-urlencoded; charset=UTF-8",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36",
"X-Requested-With": "XMLHttpRequest",
"Sec-Fetch-Dest": "empty",
"Sec-Fetch-Mode": "cors",
"Sec-Fetch-Site": "same-origin"
},
timeout=request_timeout,
follow_redirects=True
)
async def _initialize_session_for_endpoint(self, endpoint_type: str) -> bool:
"""
Initialize session and obtain CSRF token for specific endpoint.
Args:
endpoint_type: One of 'genel_kurul', 'temyiz_kurulu', 'daire'
Returns:
True if session initialized successfully, False otherwise
"""
page_mapping = {
'genel_kurul': self.GENEL_KURUL_PAGE,
'temyiz_kurulu': self.TEMYIZ_KURULU_PAGE,
'daire': self.DAIRE_PAGE
}
if endpoint_type not in page_mapping:
logger.error(f"Invalid endpoint type: {endpoint_type}")
return False
page_url = page_mapping[endpoint_type]
logger.info(f"Initializing session for {endpoint_type} endpoint: {page_url}")
try:
response = await self.http_client.get(page_url)
response.raise_for_status()
# Extract session cookies
for cookie_name, cookie_value in response.cookies.items():
self.session_cookies[cookie_name] = cookie_value
logger.debug(f"Stored session cookie: {cookie_name}")
# Extract CSRF token from form
soup = BeautifulSoup(response.text, 'html.parser')
csrf_input = soup.find('input', {'name': '__RequestVerificationToken'})
if csrf_input and csrf_input.get('value'):
self.csrf_tokens[endpoint_type] = csrf_input['value']
logger.info(f"Extracted CSRF token for {endpoint_type}")
return True
else:
logger.warning(f"CSRF token not found in {endpoint_type} page")
return False
except httpx.RequestError as e:
logger.error(f"HTTP error during session initialization for {endpoint_type}: {e}")
return False
except Exception as e:
logger.error(f"Error initializing session for {endpoint_type}: {e}")
return False
def _enum_to_form_value(self, enum_value: str, enum_type: str) -> str:
"""Convert enum values to form values expected by the API."""
if enum_value == "ALL":
if enum_type == "daire":
return "Tüm Daireler"
elif enum_type == "kamu_idaresi":
return "Tüm Kurumlar"
elif enum_type == "web_karar_konusu":
return "Tüm Konular"
# Apply web_karar_konusu mapping
if enum_type == "web_karar_konusu":
return WEB_KARAR_KONUSU_MAPPING.get(enum_value, enum_value)
return enum_value
def _build_datatables_params(self, start: int, length: int, draw: int = 1) -> List[Tuple[str, str]]:
"""Build standard DataTables parameters for all endpoints."""
params = [
("draw", str(draw)),
("start", str(start)),
("length", str(length)),
("search[value]", ""),
("search[regex]", "false")
]
return params
def _build_genel_kurul_form_data(self, params: GenelKurulSearchRequest, draw: int = 1) -> List[Tuple[str, str]]:
"""Build form data for Genel Kurul search request."""
form_data = self._build_datatables_params(params.start, params.length, draw)
# Add DataTables column definitions (from actual request)
column_defs = [
("columns[0][data]", "KARARNO"),
("columns[0][name]", ""),
("columns[0][searchable]", "true"),
("columns[0][orderable]", "false"),
("columns[0][search][value]", ""),
("columns[0][search][regex]", "false"),
("columns[1][data]", "KARARNO"),
("columns[1][name]", ""),
("columns[1][searchable]", "true"),
("columns[1][orderable]", "true"),
("columns[1][search][value]", ""),
("columns[1][search][regex]", "false"),
("columns[2][data]", "KARARTARIH"),
("columns[2][name]", ""),
("columns[2][searchable]", "true"),
("columns[2][orderable]", "true"),
("columns[2][search][value]", ""),
("columns[2][search][regex]", "false"),
("columns[3][data]", "KARAROZETI"),
("columns[3][name]", ""),
("columns[3][searchable]", "true"),
("columns[3][orderable]", "false"),
("columns[3][search][value]", ""),
("columns[3][search][regex]", "false"),
("columns[4][data]", ""),
("columns[4][name]", ""),
("columns[4][searchable]", "true"),
("columns[4][orderable]", "false"),
("columns[4][search][value]", ""),
("columns[4][search][regex]", "false"),
("order[0][column]", "2"),
("order[0][dir]", "desc")
]
form_data.extend(column_defs)
# Add search parameters
form_data.extend([
("KararlarGenelKurulAra.KARARNO", params.karar_no or ""),
("__Invariant[]", "KararlarGenelKurulAra.KARARNO"),
("__Invariant[]", "KararlarGenelKurulAra.KARAREK"),
("KararlarGenelKurulAra.KARAREK", params.karar_ek or ""),
("KararlarGenelKurulAra.KARARTARIHBaslangic", params.karar_tarih_baslangic or "Başlangıç Tarihi"),
("KararlarGenelKurulAra.KARARTARIHBitis", params.karar_tarih_bitis or "Bitiş Tarihi"),
("KararlarGenelKurulAra.KARARTAMAMI", params.karar_tamami or ""),
("__RequestVerificationToken", self.csrf_tokens.get('genel_kurul', ''))
])
return form_data
def _build_temyiz_kurulu_form_data(self, params: TemyizKuruluSearchRequest, draw: int = 1) -> List[Tuple[str, str]]:
"""Build form data for Temyiz Kurulu search request."""
form_data = self._build_datatables_params(params.start, params.length, draw)
# Add DataTables column definitions (from actual request)
column_defs = [
("columns[0][data]", "TEMYIZTUTANAKTARIHI"),
("columns[0][name]", ""),
("columns[0][searchable]", "true"),
("columns[0][orderable]", "false"),
("columns[0][search][value]", ""),
("columns[0][search][regex]", "false"),
("columns[1][data]", "TEMYIZTUTANAKTARIHI"),
("columns[1][name]", ""),
("columns[1][searchable]", "true"),
("columns[1][orderable]", "true"),
("columns[1][search][value]", ""),
("columns[1][search][regex]", "false"),
("columns[2][data]", "ILAMDAIRESI"),
("columns[2][name]", ""),
("columns[2][searchable]", "true"),
("columns[2][orderable]", "true"),
("columns[2][search][value]", ""),
("columns[2][search][regex]", "false"),
("columns[3][data]", "TEMYIZKARAR"),
("columns[3][name]", ""),
("columns[3][searchable]", "true"),
("columns[3][orderable]", "false"),
("columns[3][search][value]", ""),
("columns[3][search][regex]", "false"),
("columns[4][data]", ""),
("columns[4][name]", ""),
("columns[4][searchable]", "true"),
("columns[4][orderable]", "false"),
("columns[4][search][value]", ""),
("columns[4][search][regex]", "false"),
("order[0][column]", "1"),
("order[0][dir]", "desc")
]
form_data.extend(column_defs)
# Add search parameters
daire_value = self._enum_to_form_value(params.ilam_dairesi, "daire")
kamu_idaresi_value = self._enum_to_form_value(params.kamu_idaresi_turu, "kamu_idaresi")
web_karar_konusu_value = self._enum_to_form_value(params.web_karar_konusu, "web_karar_konusu")
form_data.extend([
("KararlarTemyizAra.ILAMDAIRESI", daire_value),
("KararlarTemyizAra.YILI", params.yili or ""),
("KararlarTemyizAra.KARARTRHBaslangic", params.karar_tarih_baslangic or ""),
("KararlarTemyizAra.KARARTRHBitis", params.karar_tarih_bitis or ""),
("KararlarTemyizAra.KAMUIDARESITURU", kamu_idaresi_value if kamu_idaresi_value != "Tüm Kurumlar" else ""),
("KararlarTemyizAra.ILAMNO", params.ilam_no or ""),
("KararlarTemyizAra.DOSYANO", params.dosya_no or ""),
("KararlarTemyizAra.TEMYIZTUTANAKNO", params.temyiz_tutanak_no or ""),
("__Invariant", "KararlarTemyizAra.TEMYIZTUTANAKNO"),
("KararlarTemyizAra.TEMYIZKARAR", params.temyiz_karar or ""),
("KararlarTemyizAra.WEBKARARKONUSU", web_karar_konusu_value if web_karar_konusu_value != "Tüm Konular" else ""),
("__RequestVerificationToken", self.csrf_tokens.get('temyiz_kurulu', ''))
])
return form_data
def _build_daire_form_data(self, params: DaireSearchRequest, draw: int = 1) -> List[Tuple[str, str]]:
"""Build form data for Daire search request."""
form_data = self._build_datatables_params(params.start, params.length, draw)
# Add DataTables column definitions (from actual request)
column_defs = [
("columns[0][data]", "YARGILAMADAIRESI"),
("columns[0][name]", ""),
("columns[0][searchable]", "true"),
("columns[0][orderable]", "false"),
("columns[0][search][value]", ""),
("columns[0][search][regex]", "false"),
("columns[1][data]", "KARARTRH"),
("columns[1][name]", ""),
("columns[1][searchable]", "true"),
("columns[1][orderable]", "true"),
("columns[1][search][value]", ""),
("columns[1][search][regex]", "false"),
("columns[2][data]", "KARARNO"),
("columns[2][name]", ""),
("columns[2][searchable]", "true"),
("columns[2][orderable]", "true"),
("columns[2][search][value]", ""),
("columns[2][search][regex]", "false"),
("columns[3][data]", "YARGILAMADAIRESI"),
("columns[3][name]", ""),
("columns[3][searchable]", "true"),
("columns[3][orderable]", "true"),
("columns[3][search][value]", ""),
("columns[3][search][regex]", "false"),
("columns[4][data]", "WEBKARARMETNI"),
("columns[4][name]", ""),
("columns[4][searchable]", "true"),
("columns[4][orderable]", "false"),
("columns[4][search][value]", ""),
("columns[4][search][regex]", "false"),
("columns[5][data]", ""),
("columns[5][name]", ""),
("columns[5][searchable]", "true"),
("columns[5][orderable]", "false"),
("columns[5][search][value]", ""),
("columns[5][search][regex]", "false"),
("order[0][column]", "2"),
("order[0][dir]", "desc")
]
form_data.extend(column_defs)
# Add search parameters
daire_value = self._enum_to_form_value(params.yargilama_dairesi, "daire")
kamu_idaresi_value = self._enum_to_form_value(params.kamu_idaresi_turu, "kamu_idaresi")
web_karar_konusu_value = self._enum_to_form_value(params.web_karar_konusu, "web_karar_konusu")
form_data.extend([
("KararlarDaireAra.YARGILAMADAIRESI", daire_value),
("KararlarDaireAra.KARARTRHBaslangic", params.karar_tarih_baslangic or ""),
("KararlarDaireAra.KARARTRHBitis", params.karar_tarih_bitis or ""),
("KararlarDaireAra.ILAMNO", params.ilam_no or ""),
("KararlarDaireAra.KAMUIDARESITURU", kamu_idaresi_value if kamu_idaresi_value != "Tüm Kurumlar" else ""),
("KararlarDaireAra.HESAPYILI", params.hesap_yili or ""),
("KararlarDaireAra.WEBKARARKONUSU", web_karar_konusu_value if web_karar_konusu_value != "Tüm Konular" else ""),
("KararlarDaireAra.WEBKARARMETNI", params.web_karar_metni or ""),
("__RequestVerificationToken", self.csrf_tokens.get('daire', ''))
])
return form_data
async def search_genel_kurul_decisions(self, params: GenelKurulSearchRequest) -> GenelKurulSearchResponse:
"""
Search Sayıştay Genel Kurul (General Assembly) decisions.
Args:
params: Search parameters for Genel Kurul decisions
Returns:
GenelKurulSearchResponse with matching decisions
"""
# Initialize session if needed
if 'genel_kurul' not in self.csrf_tokens:
if not await self._initialize_session_for_endpoint('genel_kurul'):
raise Exception("Failed to initialize session for Genel Kurul endpoint")
form_data = self._build_genel_kurul_form_data(params)
encoded_data = urlencode(form_data, encoding='utf-8')
logger.info(f"Searching Genel Kurul decisions with parameters: {params.model_dump(exclude_none=True)}")
try:
# Update headers with cookies
headers = self.http_client.headers.copy()
if self.session_cookies:
cookie_header = "; ".join([f"{k}={v}" for k, v in self.session_cookies.items()])
headers["Cookie"] = cookie_header
response = await self.http_client.post(
self.GENEL_KURUL_ENDPOINT,
data=encoded_data,
headers=headers
)
response.raise_for_status()
response_json = response.json()
# Parse response
decisions = []
for item in response_json.get('data', []):
decisions.append(GenelKurulDecision(
id=item['Id'],
karar_no=item['KARARNO'],
karar_tarih=item['KARARTARIH'],
karar_ozeti=item['KARAROZETI']
))
return GenelKurulSearchResponse(
decisions=decisions,
total_records=response_json.get('recordsTotal', 0),
total_filtered=response_json.get('recordsFiltered', 0),
draw=response_json.get('draw', 1)
)
except httpx.RequestError as e:
logger.error(f"HTTP error during Genel Kurul search: {e}")
raise
except Exception as e:
logger.error(f"Error processing Genel Kurul search: {e}")
raise
async def search_temyiz_kurulu_decisions(self, params: TemyizKuruluSearchRequest) -> TemyizKuruluSearchResponse:
"""
Search Sayıştay Temyiz Kurulu (Appeals Board) decisions.
Args:
params: Search parameters for Temyiz Kurulu decisions
Returns:
TemyizKuruluSearchResponse with matching decisions
"""
# Initialize session if needed
if 'temyiz_kurulu' not in self.csrf_tokens:
if not await self._initialize_session_for_endpoint('temyiz_kurulu'):
raise Exception("Failed to initialize session for Temyiz Kurulu endpoint")
form_data = self._build_temyiz_kurulu_form_data(params)
encoded_data = urlencode(form_data, encoding='utf-8')
logger.info(f"Searching Temyiz Kurulu decisions with parameters: {params.model_dump(exclude_none=True)}")
try:
# Update headers with cookies
headers = self.http_client.headers.copy()
if self.session_cookies:
cookie_header = "; ".join([f"{k}={v}" for k, v in self.session_cookies.items()])
headers["Cookie"] = cookie_header
response = await self.http_client.post(
self.TEMYIZ_KURULU_ENDPOINT,
data=encoded_data,
headers=headers
)
response.raise_for_status()
response_json = response.json()
# Parse response
decisions = []
for item in response_json.get('data', []):
decisions.append(TemyizKuruluDecision(
id=item['Id'],
temyiz_tutanak_tarihi=item['TEMYIZTUTANAKTARIHI'],
ilam_dairesi=item['ILAMDAIRESI'],
temyiz_karar=item['TEMYIZKARAR']
))
return TemyizKuruluSearchResponse(
decisions=decisions,
total_records=response_json.get('recordsTotal', 0),
total_filtered=response_json.get('recordsFiltered', 0),
draw=response_json.get('draw', 1)
)
except httpx.RequestError as e:
logger.error(f"HTTP error during Temyiz Kurulu search: {e}")
raise
except Exception as e:
logger.error(f"Error processing Temyiz Kurulu search: {e}")
raise
async def search_daire_decisions(self, params: DaireSearchRequest) -> DaireSearchResponse:
"""
Search Sayıştay Daire (Chamber) decisions.
Args:
params: Search parameters for Daire decisions
Returns:
DaireSearchResponse with matching decisions
"""
# Initialize session if needed
if 'daire' not in self.csrf_tokens:
if not await self._initialize_session_for_endpoint('daire'):
raise Exception("Failed to initialize session for Daire endpoint")
form_data = self._build_daire_form_data(params)
encoded_data = urlencode(form_data, encoding='utf-8')
logger.info(f"Searching Daire decisions with parameters: {params.model_dump(exclude_none=True)}")
try:
# Update headers with cookies
headers = self.http_client.headers.copy()
if self.session_cookies:
cookie_header = "; ".join([f"{k}={v}" for k, v in self.session_cookies.items()])
headers["Cookie"] = cookie_header
response = await self.http_client.post(
self.DAIRE_ENDPOINT,
data=encoded_data,
headers=headers
)
response.raise_for_status()
response_json = response.json()
# Parse response
decisions = []
for item in response_json.get('data', []):
decisions.append(DaireDecision(
id=item['Id'],
yargilama_dairesi=item['YARGILAMADAIRESI'],
karar_tarih=item['KARARTRH'],
karar_no=item['KARARNO'],
ilam_no=item.get('ILAMNO'), # Use get() to handle None values
madde_no=item['MADDENO'],
kamu_idaresi_turu=item['KAMUIDARESITURU'],
hesap_yili=item['HESAPYILI'],
web_karar_konusu=item['WEBKARARKONUSU'],
web_karar_metni=item['WEBKARARMETNI']
))
return DaireSearchResponse(
decisions=decisions,
total_records=response_json.get('recordsTotal', 0),
total_filtered=response_json.get('recordsFiltered', 0),
draw=response_json.get('draw', 1)
)
except httpx.RequestError as e:
logger.error(f"HTTP error during Daire search: {e}")
raise
except Exception as e:
logger.error(f"Error processing Daire search: {e}")
raise
def _convert_html_to_markdown(self, html_content: str) -> Optional[str]:
"""Convert HTML content to Markdown using MarkItDown with BytesIO to avoid filename length issues."""
if not html_content:
return None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_content.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
result = md_converter.convert(html_stream)
markdown_content = result.text_content
logger.info("Successfully converted HTML to Markdown")
return markdown_content
except Exception as e:
logger.error(f"Error converting HTML to Markdown: {e}")
return f"Error converting HTML content: {str(e)}"
async def get_document_as_markdown(self, decision_id: str, decision_type: str) -> SayistayDocumentMarkdown:
"""
Retrieve full text of a Sayıştay decision and convert to Markdown.
Args:
decision_id: Unique decision identifier
decision_type: Type of decision ('genel_kurul', 'temyiz_kurulu', 'daire')
Returns:
SayistayDocumentMarkdown with converted content
"""
logger.info(f"Retrieving document for {decision_type} decision ID: {decision_id}")
# Validate decision_id
if not decision_id or not decision_id.strip():
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url="",
markdown_content=None,
error_message="Decision ID cannot be empty"
)
# Map decision type to URL path
url_path_mapping = {
'genel_kurul': 'KararlarGenelKurul',
'temyiz_kurulu': 'KararlarTemyiz',
'daire': 'KararlarDaire'
}
if decision_type not in url_path_mapping:
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url="",
markdown_content=None,
error_message=f"Invalid decision type: {decision_type}. Must be one of: {list(url_path_mapping.keys())}"
)
# Build document URL
url_path = url_path_mapping[decision_type]
document_url = f"{self.BASE_URL}/{url_path}/Detay/{decision_id}/"
try:
# Make HTTP GET request to document URL
headers = {
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36",
"Sec-Fetch-Dest": "document",
"Sec-Fetch-Mode": "navigate",
"Sec-Fetch-Site": "same-origin"
}
# Include session cookies if available
if self.session_cookies:
cookie_header = "; ".join([f"{k}={v}" for k, v in self.session_cookies.items()])
headers["Cookie"] = cookie_header
response = await self.http_client.get(document_url, headers=headers)
response.raise_for_status()
html_content = response.text
if not html_content or not html_content.strip():
logger.warning(f"Received empty HTML content from {document_url}")
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=None,
error_message="Document content is empty"
)
# Convert HTML to Markdown using existing method
markdown_content = self._convert_html_to_markdown(html_content)
if markdown_content and "Error converting HTML content" not in markdown_content:
logger.info(f"Successfully retrieved and converted document {decision_id} to Markdown")
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=markdown_content,
retrieval_date=None # Could add datetime.now().isoformat() if needed
)
else:
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=None,
error_message=f"Failed to convert HTML to Markdown: {markdown_content}"
)
except httpx.HTTPStatusError as e:
error_msg = f"HTTP error {e.response.status_code} when fetching document: {e}"
logger.error(f"HTTP error fetching document {decision_id}: {error_msg}")
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=None,
error_message=error_msg
)
except httpx.RequestError as e:
error_msg = f"Network error when fetching document: {e}"
logger.error(f"Network error fetching document {decision_id}: {error_msg}")
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=None,
error_message=error_msg
)
except Exception as e:
error_msg = f"Unexpected error when fetching document: {e}"
logger.error(f"Unexpected error fetching document {decision_id}: {error_msg}")
return SayistayDocumentMarkdown(
decision_id=decision_id,
decision_type=decision_type,
source_url=document_url,
markdown_content=None,
error_message=error_msg
)
async def close_client_session(self):
"""Close HTTP client session."""
if hasattr(self, 'http_client') and self.http_client and not self.http_client.is_closed:
await self.http_client.aclose()
logger.info("SayistayApiClient: HTTP client session closed.")
@@ -1,61 +0,0 @@
# sayistay_mcp_module/enums.py
from typing import Literal
# Chamber/Daire options for Temyiz Kurulu and Daire endpoints (1-8 + All)
DaireEnum = Literal[
"ALL", # All chambers/departments
"1", # 1. Daire
"2", # 2. Daire
"3", # 3. Daire
"4", # 4. Daire
"5", # 5. Daire
"6", # 6. Daire
"7", # 7. Daire
"8" # 8. Daire
]
# Public Administration Types (Kamu İdaresi Türü)
KamuIdaresiTuruEnum = Literal[
"ALL", # All institutions
"Genel Bütçe Kapsamındaki İdareler", # General Budget Administrations
"Yüksek Öğretim Kurumları", # Higher Education Institutions
"Diğer Özel Bütçeli İdareler", # Other Special Budget Administrations
"Düzenleyici ve Denetleyici Kurumlar", # Regulatory and Supervisory Institutions
"Sosyal Güvenlik Kurumları", # Social Security Institutions
"Özel İdareler", # Special Administrations
"Belediyeler ve Bağlı İdareler", # Municipalities and Affiliated Administrations
"Diğer" # Other
]
# Decision Subject Categories (Web Karar Konusu) - Shortened for token efficiency
WebKararKonusuEnum = Literal[
"ALL", # All subjects
"Harcırah Mevzuatı", # Travel Allowance Legislation
"İhale Mevzuatı", # Procurement Legislation
"İş Mevzuatı", # Labor Legislation
"Personel Mevzuatı", # Personnel Legislation
"Sorumluluk ve Yargılama Usulleri", # Liability and Trial Procedures
"Vergi Resmi Harç ve Diğer Gelirler", # Tax, Official Fee and Other Revenue
"Çeşitli Konular" # Various Topics
]
# Mapping from shortened enum values to full API values
WEB_KARAR_KONUSU_MAPPING = {
"ALL": "ALL",
"Harcırah Mevzuatı": "Harcırah Mevzuatı ile İlgili Kararlar",
"İhale Mevzuatı": "İhale Mevzuatı ile İlgili Kararlar",
"İş Mevzuatı": "İş Mevzuatı ile İlgili Kararlar",
"Personel Mevzuatı": "Personel Mevzuatı ile İlgili Kararlar",
"Sorumluluk ve Yargılama Usulleri": "Sorumluluk ve Yargılama Usulleri ile İlgili Kararlar",
"Vergi Resmi Harç ve Diğer Gelirler": "Vergi Resmi Harç ve Diğer Gelirlerle İlgili Kararlar",
"Çeşitli Konular": "Çeşitli Konuları İlgilendiren Kararlar"
}
# Year ranges for different endpoints
GENEL_KURUL_YEARS = [str(year) for year in range(2006, 2025)] # 2006-2024
TEMYIZ_KURULU_YEARS = [str(year) for year in range(1993, 2023)] # 1993-2022
DAIRE_YEARS = [str(year) for year in range(2012, 2026)] # 2012-2025
# Account years for Temyiz Kurulu and Daire endpoints
HESAP_YILLARI = [str(year) for year in range(1993, 2024)] # 1993-2023
@@ -1,220 +0,0 @@
# sayistay_mcp_module/models.py
from pydantic import BaseModel, Field
from typing import Optional, List, Union, Dict, Any, Literal
from enum import Enum
from .enums import DaireEnum, KamuIdaresiTuruEnum, WebKararKonusuEnum
# --- Unified Enums ---
class SayistayDecisionTypeEnum(str, Enum):
GENEL_KURUL = "genel_kurul"
TEMYIZ_KURULU = "temyiz_kurulu"
DAIRE = "daire"
# ============================================================================
# Genel Kurul (General Assembly) Models
# ============================================================================
class GenelKurulSearchRequest(BaseModel):
"""
Search request for Sayıştay Genel Kurul (General Assembly) decisions.
Genel Kurul decisions are precedent-setting rulings made by the full assembly
of the Turkish Court of Accounts, typically addressing interpretation of
audit and accountability regulations.
"""
karar_no: str = Field("", description="Decision no")
karar_ek: str = Field("", description="Appendix no")
karar_tarih_baslangic: str = Field("", description="Start year (YYYY)")
karar_tarih_bitis: str = Field("", description="End year")
karar_tamami: str = Field("", description="Value")
# DataTables pagination
start: int = Field(0, description="Starting record for pagination (0-based)")
length: int = Field(10, description="Number of records per page (1-10)")
class GenelKurulDecision(BaseModel):
"""Single Genel Kurul decision entry from search results."""
id: int = Field(..., description="Unique decision ID")
karar_no: str = Field(..., description="Decision number (e.g., '5415/1')")
karar_tarih: str = Field(..., description="Decision date in DD.MM.YYYY format")
karar_ozeti: str = Field(..., description="Decision summary/abstract")
class GenelKurulSearchResponse(BaseModel):
"""Response from Genel Kurul search endpoint."""
decisions: List[GenelKurulDecision] = Field(default_factory=list, description="List of matching decisions")
total_records: int = Field(0, description="Total number of matching records")
total_filtered: int = Field(0, description="Number of records after filtering")
draw: int = Field(1, description="DataTables draw counter")
# ============================================================================
# Temyiz Kurulu (Appeals Board) Models
# ============================================================================
class TemyizKuruluSearchRequest(BaseModel):
"""
Search request for Sayıştay Temyiz Kurulu (Appeals Board) decisions.
Temyiz Kurulu reviews appeals against audit chamber decisions,
providing higher-level review of audit findings and sanctions.
"""
ilam_dairesi: DaireEnum = Field("ALL", description="Value")
yili: str = Field("", description="Value")
karar_tarih_baslangic: str = Field("", description="Value")
karar_tarih_bitis: str = Field("", description="End year")
kamu_idaresi_turu: KamuIdaresiTuruEnum = Field("ALL", description="Value")
ilam_no: str = Field("", description="Audit report number (İlam No, max 50 chars)")
dosya_no: str = Field("", description="File number for the case")
temyiz_tutanak_no: str = Field("", description="Appeals board meeting minutes number")
temyiz_karar: str = Field("", description="Value")
web_karar_konusu: WebKararKonusuEnum = Field("ALL", description="Value")
# DataTables pagination
start: int = Field(0, description="Starting record for pagination (0-based)")
length: int = Field(10, description="Number of records per page (1-10)")
class TemyizKuruluDecision(BaseModel):
"""Single Temyiz Kurulu decision entry from search results."""
id: int = Field(..., description="Unique decision ID")
temyiz_tutanak_tarihi: str = Field(..., description="Appeals board meeting date in DD.MM.YYYY format")
ilam_dairesi: int = Field(..., description="Chamber number (1-8)")
temyiz_karar: str = Field(..., description="Appeals decision summary and reasoning")
class TemyizKuruluSearchResponse(BaseModel):
"""Response from Temyiz Kurulu search endpoint."""
decisions: List[TemyizKuruluDecision] = Field(default_factory=list, description="List of matching appeals decisions")
total_records: int = Field(0, description="Total number of matching records")
total_filtered: int = Field(0, description="Number of records after filtering")
draw: int = Field(1, description="DataTables draw counter")
# ============================================================================
# Daire (Chamber) Models
# ============================================================================
class DaireSearchRequest(BaseModel):
"""
Search request for Sayıştay Daire (Chamber) decisions.
Daire decisions are first-instance audit findings and sanctions
issued by individual audit chambers before potential appeals.
"""
yargilama_dairesi: DaireEnum = Field("ALL", description="Value")
karar_tarih_baslangic: str = Field("", description="Value")
karar_tarih_bitis: str = Field("", description="End year")
ilam_no: str = Field("", description="Audit report number (İlam No, max 50 chars)")
kamu_idaresi_turu: KamuIdaresiTuruEnum = Field("ALL", description="Value")
hesap_yili: str = Field("", description="Value")
web_karar_konusu: WebKararKonusuEnum = Field("ALL", description="Value")
web_karar_metni: str = Field("", description="Value")
# DataTables pagination
start: int = Field(0, description="Starting record for pagination (0-based)")
length: int = Field(10, description="Number of records per page (1-10)")
class DaireDecision(BaseModel):
"""Single Daire decision entry from search results."""
id: int = Field(..., description="Unique decision ID")
yargilama_dairesi: int = Field(..., description="Chamber number (1-8)")
karar_tarih: str = Field(..., description="Decision date in DD.MM.YYYY format")
karar_no: str = Field(..., description="Decision number")
ilam_no: str = Field("", description="Audit report number (may be null)")
madde_no: int = Field(..., description="Article/item number within the decision")
kamu_idaresi_turu: str = Field(..., description="Public administration type")
hesap_yili: int = Field(..., description="Account year being audited")
web_karar_konusu: str = Field(..., description="Decision subject category")
web_karar_metni: str = Field(..., description="Decision text/summary")
class DaireSearchResponse(BaseModel):
"""Response from Daire search endpoint."""
decisions: List[DaireDecision] = Field(default_factory=list, description="List of matching chamber decisions")
total_records: int = Field(0, description="Total number of matching records")
total_filtered: int = Field(0, description="Number of records after filtering")
draw: int = Field(1, description="DataTables draw counter")
# ============================================================================
# Document Models
# ============================================================================
class SayistayDocumentMarkdown(BaseModel):
"""
Sayıştay decision document converted to Markdown format.
Used for retrieving full text of decisions from any of the three
decision types (Genel Kurul, Temyiz Kurulu, Daire).
"""
decision_id: str = Field(..., description="Unique decision identifier")
decision_type: str = Field(..., description="Value")
source_url: str = Field(..., description="Original URL where the document was retrieved")
markdown_content: Optional[str] = Field(None, description="Full decision text converted to Markdown format")
retrieval_date: Optional[str] = Field(None, description="Date when document was retrieved (ISO format)")
error_message: Optional[str] = Field(None, description="Error message if document retrieval failed")
# ============================================================================
# Unified Models
# ============================================================================
class SayistayUnifiedSearchRequest(BaseModel):
"""Unified search request for all Sayıştay decision types."""
decision_type: Literal["genel_kurul", "temyiz_kurulu", "daire"] = Field(..., description="Decision type: genel_kurul, temyiz_kurulu, or daire")
# Common pagination parameters
start: int = Field(0, ge=0, description="Starting record for pagination (0-based)")
length: int = Field(10, ge=1, le=100, description="Number of records per page (1-100)")
# Common search parameters
karar_tarih_baslangic: str = Field("", description="Start date (DD.MM.YYYY format)")
karar_tarih_bitis: str = Field("", description="End date (DD.MM.YYYY format)")
kamu_idaresi_turu: KamuIdaresiTuruEnum = Field("ALL", description="Public administration type filter")
ilam_no: str = Field("", description="Audit report number (İlam No, max 50 chars)")
web_karar_konusu: WebKararKonusuEnum = Field("ALL", description="Decision subject category filter")
# Genel Kurul specific parameters (ignored for other types)
karar_no: str = Field("", description="Decision number (genel_kurul only)")
karar_ek: str = Field("", description="Decision appendix number (genel_kurul only)")
karar_tamami: str = Field("", description="Full text search (genel_kurul only)")
# Temyiz Kurulu specific parameters (ignored for other types)
ilam_dairesi: DaireEnum = Field("ALL", description="Audit chamber selection (temyiz_kurulu only)")
yili: str = Field("", description="Year (YYYY format, temyiz_kurulu only)")
dosya_no: str = Field("", description="File number (temyiz_kurulu only)")
temyiz_tutanak_no: str = Field("", description="Appeals board meeting minutes number (temyiz_kurulu only)")
temyiz_karar: str = Field("", description="Appeals decision text search (temyiz_kurulu only)")
# Daire specific parameters (ignored for other types)
yargilama_dairesi: DaireEnum = Field("ALL", description="Chamber selection (daire only)")
hesap_yili: str = Field("", description="Account year (daire only)")
web_karar_metni: str = Field("", description="Decision text search (daire only)")
class SayistayUnifiedSearchResult(BaseModel):
"""Unified search result containing decisions from any Sayıştay decision type."""
decision_type: Literal["genel_kurul", "temyiz_kurulu", "daire"] = Field(..., description="Type of decisions returned")
decisions: List[Dict[str, Any]] = Field(default_factory=list, description="Decision list (structure varies by type)")
total_records: int = Field(0, description="Total number of records found")
total_filtered: int = Field(0, description="Number of records after filtering")
draw: int = Field(1, description="DataTables draw counter")
class SayistayUnifiedDocumentMarkdown(BaseModel):
"""Unified document model for all Sayıştay decision types."""
decision_type: Literal["genel_kurul", "temyiz_kurulu", "daire"] = Field(..., description="Type of document")
decision_id: str = Field(..., description="Decision ID")
source_url: str = Field(..., description="Source URL of the document")
document_data: Dict[str, Any] = Field(default_factory=dict, description="Document content and metadata")
markdown_content: Optional[str] = Field(None, description="Markdown content")
error_message: Optional[str] = Field(None, description="Error message if retrieval failed")
@@ -1,133 +0,0 @@
# sayistay_mcp_module/unified_client.py
# Unified client for all three Sayıştay decision types
import logging
from typing import Optional, Dict, Any
from urllib.parse import urlparse
from .models import (
SayistayUnifiedSearchRequest,
SayistayUnifiedSearchResult,
SayistayUnifiedDocumentMarkdown,
GenelKurulSearchRequest,
TemyizKuruluSearchRequest,
DaireSearchRequest
)
from .client import SayistayApiClient
logger = logging.getLogger(__name__)
class SayistayUnifiedClient:
"""Unified client that handles all three Sayıştay decision types."""
def __init__(self, request_timeout: float = 60.0):
self.client = SayistayApiClient(request_timeout)
async def search_unified(self, params: SayistayUnifiedSearchRequest) -> SayistayUnifiedSearchResult:
"""Unified search that routes to appropriate search method based on decision_type."""
if params.decision_type == "genel_kurul":
# Convert to genel kurul request
genel_kurul_params = GenelKurulSearchRequest(
karar_no=params.karar_no,
karar_ek=params.karar_ek,
karar_tarih_baslangic=params.karar_tarih_baslangic,
karar_tarih_bitis=params.karar_tarih_bitis,
karar_tamami=params.karar_tamami,
start=params.start,
length=params.length
)
result = await self.client.search_genel_kurul_decisions(genel_kurul_params)
# Convert to unified format
decisions_list = [decision.model_dump() for decision in result.decisions]
return SayistayUnifiedSearchResult(
decision_type="genel_kurul",
decisions=decisions_list,
total_records=result.total_records,
total_filtered=result.total_filtered,
draw=result.draw
)
elif params.decision_type == "temyiz_kurulu":
# Convert to temyiz kurulu request
temyiz_params = TemyizKuruluSearchRequest(
ilam_dairesi=params.ilam_dairesi,
yili=params.yili,
karar_tarih_baslangic=params.karar_tarih_baslangic,
karar_tarih_bitis=params.karar_tarih_bitis,
kamu_idaresi_turu=params.kamu_idaresi_turu,
ilam_no=params.ilam_no,
dosya_no=params.dosya_no,
temyiz_tutanak_no=params.temyiz_tutanak_no,
temyiz_karar=params.temyiz_karar,
web_karar_konusu=params.web_karar_konusu,
start=params.start,
length=params.length
)
result = await self.client.search_temyiz_kurulu_decisions(temyiz_params)
# Convert to unified format
decisions_list = [decision.model_dump() for decision in result.decisions]
return SayistayUnifiedSearchResult(
decision_type="temyiz_kurulu",
decisions=decisions_list,
total_records=result.total_records,
total_filtered=result.total_filtered,
draw=result.draw
)
elif params.decision_type == "daire":
# Convert to daire request
daire_params = DaireSearchRequest(
yargilama_dairesi=params.yargilama_dairesi,
karar_tarih_baslangic=params.karar_tarih_baslangic,
karar_tarih_bitis=params.karar_tarih_bitis,
ilam_no=params.ilam_no,
kamu_idaresi_turu=params.kamu_idaresi_turu,
hesap_yili=params.hesap_yili,
web_karar_konusu=params.web_karar_konusu,
web_karar_metni=params.web_karar_metni,
start=params.start,
length=params.length
)
result = await self.client.search_daire_decisions(daire_params)
# Convert to unified format
decisions_list = [decision.model_dump() for decision in result.decisions]
return SayistayUnifiedSearchResult(
decision_type="daire",
decisions=decisions_list,
total_records=result.total_records,
total_filtered=result.total_filtered,
draw=result.draw
)
else:
raise ValueError(f"Unsupported decision type: {params.decision_type}")
async def get_document_unified(self, decision_id: str, decision_type: str) -> SayistayUnifiedDocumentMarkdown:
"""Unified document retrieval for all Sayıştay decision types."""
# Use existing client method (decision_type is already a string)
result = await self.client.get_document_as_markdown(decision_id, decision_type)
return SayistayUnifiedDocumentMarkdown(
decision_type=decision_type,
decision_id=result.decision_id,
source_url=result.source_url,
document_data=result.model_dump(),
markdown_content=result.markdown_content,
error_message=result.error_message
)
async def close_client_session(self):
"""Close the underlying client session."""
if hasattr(self.client, 'close_client_session'):
await self.client.close_client_session()
@@ -1,159 +0,0 @@
"""
Starlette integration example for Yargı MCP Server
This module demonstrates how to integrate the Yargı MCP server
with a Starlette application, including authentication middleware
and custom routing.
Usage:
uvicorn starlette_app:app --host 0.0.0.0 --port 8000
"""
import os
from starlette.applications import Starlette
from starlette.routing import Mount, Route
from starlette.requests import Request
from starlette.responses import JSONResponse, PlainTextResponse, RedirectResponse
from starlette.middleware import Middleware
from starlette.middleware.cors import CORSMiddleware
from starlette.middleware.authentication import AuthenticationMiddleware
from starlette.authentication import (
AuthenticationBackend, AuthCredentials, SimpleUser, AuthenticationError
)
# Import the main MCP app
from mcp_server_main import app as mcp_server
# Simple token authentication backend
class TokenAuthBackend(AuthenticationBackend):
async def authenticate(self, request):
auth_header = request.headers.get("Authorization")
expected_token = os.getenv("API_TOKEN")
# Skip auth for health check and public endpoints
if request.url.path in ["/health", "/", "/login"]:
return None
if not expected_token:
# No token configured, allow all
return AuthCredentials(["authenticated"]), SimpleUser("anonymous")
if not auth_header:
raise AuthenticationError("Authorization header required")
try:
scheme, token = auth_header.split()
if scheme.lower() != "bearer":
raise AuthenticationError("Invalid authentication scheme")
if token != expected_token:
raise AuthenticationError("Invalid token")
return AuthCredentials(["authenticated"]), SimpleUser("user")
except ValueError:
raise AuthenticationError("Invalid authorization header format")
# Homepage
async def homepage(request: Request):
return JSONResponse({
"service": "Yargı MCP Server",
"version": "0.1.0",
"endpoints": {
"mcp": "/mcp-server/mcp/",
"api": "/api/",
"health": "/health"
}
})
# API info endpoint
async def api_info(request: Request):
if not request.user.is_authenticated:
return JSONResponse({"error": "Authentication required"}, status_code=401)
return JSONResponse({
"authenticated_as": request.user.display_name,
"available_tools": len(mcp_server._tool_manager._tools),
"databases": [
"Yargıtay", "Danıştay", "Emsal", "Uyuşmazlık",
"Anayasa", "KIK", "Rekabet", "Bedesten"
]
})
# Health check
async def health_check(request: Request):
return JSONResponse({
"status": "healthy",
"service": "Yargı MCP Server"
})
# Login example (returns token for demo)
async def login(request: Request):
token = os.getenv("API_TOKEN", "demo-token")
return JSONResponse({
"message": "Use this token in Authorization header",
"example": f"Authorization: Bearer {token}",
"note": "Set API_TOKEN environment variable to change token"
})
# Create MCP ASGI app
mcp_app = mcp_server.http_app(path='/mcp')
# Configure middleware
middleware = [
Middleware(
CORSMiddleware,
allow_origins=os.getenv("ALLOWED_ORIGINS", "*").split(","),
allow_credentials=True,
allow_methods=["*"],
allow_headers=["*"],
),
Middleware(AuthenticationMiddleware, backend=TokenAuthBackend()),
]
# Create routes
routes = [
Route("/", homepage),
Route("/health", health_check),
Route("/login", login),
Route("/api/info", api_info),
Mount("/mcp-server", app=mcp_app),
]
# Create Starlette app
app = Starlette(
routes=routes,
middleware=middleware,
lifespan=mcp_app.lifespan
)
# Nested mount example
def create_nested_app():
"""Example of nested mounting for complex routing structures"""
# Create inner app with MCP
inner_app = Starlette(
routes=[Mount("/services", app=mcp_app)],
middleware=middleware
)
# Create outer app
outer_app = Starlette(
routes=[
Route("/", homepage),
Mount("/v1", app=inner_app),
],
lifespan=mcp_app.lifespan
)
# MCP would be available at /v1/services/mcp/
return outer_app
# Export both apps
nested_app = create_nested_app()
if __name__ == "__main__":
import uvicorn
print("Starting Starlette app with authentication...")
print("Set API_TOKEN environment variable to enable authentication")
print("Example: API_TOKEN=secret-token python starlette_app.py")
uvicorn.run(app, host="0.0.0.0", port=8000)
@@ -1,25 +0,0 @@
import os, stripe
from clerk_backend_api import Clerk # Clerk backend SDK
from fastapi import APIRouter, Request, HTTPException
router = APIRouter()
stripe.api_key = os.getenv("STRIPE_SECRET")
clerk = Clerk(bearer_auth=os.getenv("CLERK_SECRET_KEY"))
@router.post("/stripe/webhook")
async def stripe_hook(req: Request):
payload, sig = await req.body(), req.headers["stripe-signature"]
try:
event = stripe.Webhook.construct_event( # Stripe-recommended verify
payload, sig, os.getenv("STRIPE_WEBHOOK_SECRET"))
except stripe.error.SignatureVerificationError:
raise HTTPException(400, "Bad sig")
if event["type"] == "customer.subscription.updated":
item = event["data"]["object"]["items"]["data"][0]
plan = item["price"]["nickname"] # "Pro", "Enterprise"…
userID = event["data"]["object"]["metadata"]["clerk_user_id"]
clerk.users.update_user_metadata( # merge into unsafe_metadata
userID, unsafe_metadata={"plan": plan})
return {"ok": True}
@@ -1,240 +0,0 @@
# uyusmazlik_mcp_module/client.py
import httpx
import aiohttp
from bs4 import BeautifulSoup
from typing import Dict, Any, List, Optional, Union, Tuple
import logging
import html
import re
import io
from markitdown import MarkItDown
from urllib.parse import urljoin, urlencode # urlencode for aiohttp form data
from .models import (
UyusmazlikSearchRequest,
UyusmazlikApiDecisionEntry,
UyusmazlikSearchResponse,
UyusmazlikDocumentMarkdown,
UyusmazlikBolumEnum,
UyusmazlikTuruEnum,
UyusmazlikKararSonucuEnum
)
logger = logging.getLogger(__name__)
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
# --- Mappings from user-friendly Enum values to API IDs ---
BOLUM_ENUM_TO_ID_MAP = {
UyusmazlikBolumEnum.CEZA_BOLUMU: "f6b74320-f2d7-4209-ad6e-c6df180d4e7c",
UyusmazlikBolumEnum.GENEL_KURUL_KARARLARI: "e4ca658d-a75a-4719-b866-b2d2f1c3b1d9",
UyusmazlikBolumEnum.HUKUK_BOLUMU: "96b26fc4-ef8e-4a4f-a9cc-a3de89952aa1",
UyusmazlikBolumEnum.TUMU: "", # Represents "...Seçiniz..." or all - empty string for API
"ALL": "" # Also map the new "ALL" literal to empty string for backward compatibility
}
UYUSMAZLIK_TURU_ENUM_TO_ID_MAP = {
UyusmazlikTuruEnum.GOREV_UYUSMAZLIGI: "7b1e2cd3-8f09-418a-921c-bbe501e1740c",
UyusmazlikTuruEnum.HUKUM_UYUSMAZLIGI: "19b88402-172b-4c1d-8339-595c942a89f5",
UyusmazlikTuruEnum.TUMU: "", # Represents "...Seçiniz..." or all - empty string for API
"ALL": "" # Also map the new "ALL" literal to empty string for backward compatibility
}
KARAR_SONUCU_ENUM_TO_ID_MAP = {
# These IDs are from the form HTML provided by the user
UyusmazlikKararSonucuEnum.HUKUM_UYUSMAZLIGI_OLMADIGINA_DAIR: "6f47d87f-dcb5-412e-9878-000385dba1d9",
UyusmazlikKararSonucuEnum.HUKUM_UYUSMAZLIGI_OLDUGUNA_DAIR: "5a01742a-c440-4c4a-ba1f-da20837cffed",
# Add all other 'Karar Sonucu' enum members and their corresponding GUIDs
# by inspecting the 'KararSonucuList' checkboxes in the provided form HTML.
}
# --- End Mappings ---
class UyusmazlikApiClient:
BASE_URL = "https://kararlar.uyusmazlik.gov.tr"
SEARCH_ENDPOINT = "/Arama/Search"
# Individual documents are fetched by their full URLs obtained from search results.
def __init__(self, request_timeout: float = 30.0):
self.request_timeout = request_timeout # Store timeout for aiohttp and httpx
# Headers for aiohttp search. httpx for docs will create its own.
self.default_aiohttp_search_headers = {
"Accept": "*/*", # Mimicking browser headers provided by user
"Accept-Encoding": "gzip, deflate, br, zstd",
"Accept-Language": "tr-TR,tr;q=0.9,en-US;q=0.8,en;q=0.7",
"X-Requested-With": "XMLHttpRequest",
"Origin": self.BASE_URL,
"Referer": self.BASE_URL + "/",
}
async def search_decisions(
self,
params: UyusmazlikSearchRequest
) -> UyusmazlikSearchResponse:
bolum_id_for_api = BOLUM_ENUM_TO_ID_MAP.get(params.bolum, "")
uyusmazlik_id_for_api = UYUSMAZLIK_TURU_ENUM_TO_ID_MAP.get(params.uyusmazlik_turu, "")
form_data_list: List[Tuple[str, str]] = []
def add_to_form_data(key: str, value: Optional[str]):
# API expects empty strings for omitted optional fields based on user payload example
form_data_list.append((key, value or ""))
add_to_form_data("BolumId", bolum_id_for_api)
add_to_form_data("UyusmazlikId", uyusmazlik_id_for_api)
if params.karar_sonuclari:
for enum_member in params.karar_sonuclari:
api_id = KARAR_SONUCU_ENUM_TO_ID_MAP.get(enum_member)
if api_id: # Only add if a valid ID is found
form_data_list.append(('KararSonucuList', api_id))
add_to_form_data("EsasYil", params.esas_yil)
add_to_form_data("EsasSayisi", params.esas_sayisi)
add_to_form_data("KararYil", params.karar_yil)
add_to_form_data("KararSayisi", params.karar_sayisi)
add_to_form_data("KanunNo", params.kanun_no)
add_to_form_data("KararDateBegin", params.karar_date_begin)
add_to_form_data("KararDateEnd", params.karar_date_end)
add_to_form_data("ResmiGazeteSayi", params.resmi_gazete_sayi)
add_to_form_data("ResmiGazeteDate", params.resmi_gazete_date)
add_to_form_data("Icerik", params.icerik)
add_to_form_data("Tumce", params.tumce)
add_to_form_data("WildCard", params.wild_card)
add_to_form_data("Hepsi", params.hepsi)
add_to_form_data("Herhangibirisi", params.herhangi_birisi)
add_to_form_data("NotHepsi", params.not_hepsi)
# X-Requested-With is handled by default_aiohttp_search_headers
search_url = urljoin(self.BASE_URL, self.SEARCH_ENDPOINT)
# For aiohttp, data for application/x-www-form-urlencoded should be a dict or str.
# Using urlencode for list of tuples.
encoded_form_payload = urlencode(form_data_list, encoding='UTF-8')
logger.info(f"UyusmazlikApiClient (aiohttp): Performing search to {search_url} with form_data: {encoded_form_payload}")
html_content = ""
aiohttp_headers = self.default_aiohttp_search_headers.copy()
aiohttp_headers["Content-Type"] = "application/x-www-form-urlencoded; charset=UTF-8"
try:
# Create a new session for each call for simplicity with aiohttp here
async with aiohttp.ClientSession(headers=aiohttp_headers) as session:
async with session.post(search_url, data=encoded_form_payload, timeout=self.request_timeout) as response:
response.raise_for_status() # Raises ClientResponseError for 400-599
html_content = await response.text(encoding='utf-8') # Ensure correct encoding
logger.debug("UyusmazlikApiClient (aiohttp): Received HTML response for search.")
except aiohttp.ClientError as e:
logger.error(f"UyusmazlikApiClient (aiohttp): HTTP client error during search: {e}")
raise # Re-raise to be handled by the MCP tool
except Exception as e:
logger.error(f"UyusmazlikApiClient (aiohttp): Error processing search request: {e}")
raise
# --- HTML Parsing (remains the same as previous version) ---
soup = BeautifulSoup(html_content, 'html.parser')
total_records_text_div = soup.find("div", class_="pull-right label label-important")
total_records = None
if total_records_text_div:
match_records = re.search(r'(\d+)\s*adet kayıt bulundu', total_records_text_div.get_text(strip=True))
if match_records:
total_records = int(match_records.group(1))
result_table = soup.find("table", class_="table-hover")
processed_decisions: List[UyusmazlikApiDecisionEntry] = []
if result_table:
rows = result_table.find_all("tr")
if len(rows) > 1: # Skip header row
for row in rows[1:]:
cols = row.find_all('td')
if len(cols) >= 5:
try:
popover_div = cols[0].find("div", attrs={"data-rel": "popover"})
popover_content_raw = popover_div["data-content"] if popover_div and popover_div.has_attr("data-content") else None
link_tag = cols[0].find('a')
doc_relative_url = link_tag['href'] if link_tag and link_tag.has_attr('href') else None
if not doc_relative_url: continue
document_url_str = urljoin(self.BASE_URL, doc_relative_url)
pdf_link_tag = cols[5].find('a', href=re.compile(r'\.pdf$', re.IGNORECASE)) if len(cols) > 5 else None
pdf_url_str = urljoin(self.BASE_URL, pdf_link_tag['href']) if pdf_link_tag and pdf_link_tag.has_attr('href') else None
decision_data_parsed = {
"karar_sayisi": cols[0].get_text(strip=True),
"esas_sayisi": cols[1].get_text(strip=True),
"bolum": cols[2].get_text(strip=True),
"uyusmazlik_konusu": cols[3].get_text(strip=True),
"karar_sonucu": cols[4].get_text(strip=True),
"popover_content": html.unescape(popover_content_raw) if popover_content_raw else None,
"document_url": document_url_str,
"pdf_url": pdf_url_str
}
decision_model = UyusmazlikApiDecisionEntry(**decision_data_parsed)
processed_decisions.append(decision_model)
except Exception as e:
logger.warning(f"UyusmazlikApiClient: Could not parse decision row. Row content: {row.get_text(strip=True, separator=' | ')}, Error: {e}")
return UyusmazlikSearchResponse(
decisions=processed_decisions,
total_records_found=total_records
)
def _convert_html_to_markdown_uyusmazlik(self, full_decision_html_content: str) -> Optional[str]:
"""Converts direct HTML content (from an Uyuşmazlık decision page) to Markdown."""
if not full_decision_html_content:
return None
processed_html = html.unescape(full_decision_html_content)
# As per user request, pass the full (unescaped) HTML to MarkItDown
html_input_for_markdown = processed_html
markdown_text = None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_input_for_markdown.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_text = conversion_result.text_content
logger.info("UyusmazlikApiClient: HTML to Markdown conversion successful.")
except Exception as e:
logger.error(f"UyusmazlikApiClient: Error during MarkItDown HTML to Markdown conversion: {e}")
return markdown_text
async def get_decision_document_as_markdown(self, document_url: str) -> UyusmazlikDocumentMarkdown:
"""
Retrieves a specific Uyuşmazlık decision from its full URL and returns content as Markdown.
"""
logger.info(f"UyusmazlikApiClient (httpx for docs): Fetching Uyuşmazlık document for Markdown from URL: {document_url}")
try:
# Using a new httpx.AsyncClient instance for this GET request for simplicity
async with httpx.AsyncClient(verify=False, timeout=self.request_timeout) as doc_fetch_client:
get_response = await doc_fetch_client.get(document_url, headers={"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"})
get_response.raise_for_status()
html_content_from_api = get_response.text
if not isinstance(html_content_from_api, str) or not html_content_from_api.strip():
logger.warning(f"UyusmazlikApiClient: Received empty or non-string HTML from URL {document_url}.")
return UyusmazlikDocumentMarkdown(source_url=document_url, markdown_content=None)
markdown_content = self._convert_html_to_markdown_uyusmazlik(html_content_from_api)
return UyusmazlikDocumentMarkdown(source_url=document_url, markdown_content=markdown_content)
except httpx.RequestError as e:
logger.error(f"UyusmazlikApiClient (httpx for docs): HTTP error fetching Uyuşmazlık document from {document_url}: {e}")
raise
except Exception as e:
logger.error(f"UyusmazlikApiClient (httpx for docs): General error processing Uyuşmazlık document from {document_url}: {e}")
raise
async def close_client_session(self):
logger.info("UyusmazlikApiClient: No persistent client session from __init__ to close.")
@@ -1,86 +0,0 @@
# uyusmazlik_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl
from typing import List, Optional
from enum import Enum
# Enum definitions for user-friendly input based on the provided HTML form
class UyusmazlikBolumEnum(str, Enum):
"""User-friendly names for 'BolumId'."""
TUMU = "ALL" # Represents "...Seçiniz..." or all
CEZA_BOLUMU = "Ceza Bölümü"
GENEL_KURUL_KARARLARI = "Genel Kurul Kararları"
HUKUK_BOLUMU = "Hukuk Bölümü"
class UyusmazlikTuruEnum(str, Enum):
"""User-friendly names for 'UyusmazlikId'."""
TUMU = "ALL" # Represents "...Seçiniz..." or all
GOREV_UYUSMAZLIGI = "Görev Uyuşmazlığı"
HUKUM_UYUSMAZLIGI = "Hüküm Uyuşmazlığı"
class UyusmazlikKararSonucuEnum(str, Enum): # Based on checkbox text in the form
"""User-friendly names for 'KararSonucuList' items."""
HUKUM_UYUSMAZLIGI_OLMADIGINA_DAIR = "Hüküm Uyuşmazlığı Olmadığına Dair"
HUKUM_UYUSMAZLIGI_OLDUGUNA_DAIR = "Hüküm Uyuşmazlığı Olduğuna Dair"
# Add other "Karar Sonucu" options from the form's checkboxes as Enum members
# Example: GOREVLI_YARGI_YERI_ADLI = "Görevli Yargı Yeri Belirlenmesine Dair (Adli Yargı)"
# The client will map these enum values (which are strings) to their respective IDs.
class UyusmazlikSearchRequest(BaseModel): # This is the model the MCP tool will accept
"""Model for Uyuşmazlık Mahkemesi search request using user-friendly terms."""
icerik: Optional[str] = Field("", description="Search text")
bolum: Optional[UyusmazlikBolumEnum] = Field(
UyusmazlikBolumEnum.TUMU,
description="Department"
)
uyusmazlik_turu: Optional[UyusmazlikTuruEnum] = Field(
UyusmazlikTuruEnum.TUMU,
description="Dispute type"
)
# User provides a list of user-friendly names for Karar Sonucu
karar_sonuclari: Optional[List[UyusmazlikKararSonucuEnum]] = Field( # Changed to list of Enums
default_factory=list,
description="Decision types"
)
esas_yil: Optional[str] = Field("", description="Case year")
esas_sayisi: Optional[str] = Field("", description="Case no")
karar_yil: Optional[str] = Field("", description="Decision year")
karar_sayisi: Optional[str] = Field("", description="Decision no")
kanun_no: Optional[str] = Field("", description="Law no")
karar_date_begin: Optional[str] = Field("", description="Start date (DD.MM.YYYY)")
karar_date_end: Optional[str] = Field("", description="End date (DD.MM.YYYY)")
resmi_gazete_sayi: Optional[str] = Field("", description="Gazette no")
resmi_gazete_date: Optional[str] = Field("", description="Gazette date (DD.MM.YYYY)")
# Detailed text search fields from the "icerikDetail" section of the form
tumce: Optional[str] = Field("", description="Exact phrase")
wild_card: Optional[str] = Field("", description="Wildcard search")
hepsi: Optional[str] = Field("", description="All words")
herhangi_birisi: Optional[str] = Field("", description="Any word")
not_hepsi: Optional[str] = Field("", description="Exclude words")
class UyusmazlikApiDecisionEntry(BaseModel):
"""Model for an individual decision entry parsed from Uyuşmazlık API's HTML search response."""
karar_sayisi: Optional[str] = Field(None)
esas_sayisi: Optional[str] = Field(None)
bolum: Optional[str] = Field(None)
uyusmazlik_konusu: Optional[str] = Field(None)
karar_sonucu: Optional[str] = Field(None)
popover_content: Optional[str] = Field(None, description="Summary")
document_url: HttpUrl # Full URL to the decision document HTML page
pdf_url: Optional[HttpUrl] = Field(None, description="PDF URL")
class UyusmazlikSearchResponse(BaseModel): # This is what the MCP tool will return
"""Response model for Uyuşmazlık Mahkemesi search results for the MCP tool."""
decisions: List[UyusmazlikApiDecisionEntry]
total_records_found: Optional[int] = Field(None, description="Total number of records found for the query, if available.")
class UyusmazlikDocumentMarkdown(BaseModel):
"""Model for an Uyuşmazlık decision document, containing only Markdown content."""
source_url: HttpUrl # The URL from which the content was fetched
markdown_content: Optional[str] = Field(None, description="The decision content converted to Markdown.")
@@ -1,182 +0,0 @@
# yargitay_mcp_module/client.py
import httpx
from bs4 import BeautifulSoup # Still needed for pre-processing HTML before markitdown
from typing import Dict, Any, List, Optional
import logging
import html
import re
import io
from markitdown import MarkItDown
from .models import (
YargitayDetailedSearchRequest,
YargitayApiSearchResponse,
YargitayApiDecisionEntry,
YargitayDocumentMarkdown,
CompactYargitaySearchResult
)
logger = logging.getLogger(__name__)
# Basic logging configuration if no handlers are configured
if not logger.hasHandlers():
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
class YargitayOfficialApiClient:
"""
API Client for Yargitay's official decision search system.
Targets the detailed search endpoint (e.g., /aramadetaylist) based on user-provided payload.
"""
BASE_URL = "https://karararama.yargitay.gov.tr"
# The form action was "/detayliArama". This often maps to an API endpoint like "/aramadetaylist".
# This should be confirmed with the actual API.
DETAILED_SEARCH_ENDPOINT = "/aramadetaylist"
DOCUMENT_ENDPOINT = "/getDokuman"
def __init__(self, request_timeout: float = 60.0):
self.http_client = httpx.AsyncClient(
base_url=self.BASE_URL,
headers={
"Content-Type": "application/json; charset=UTF-8",
"Accept": "application/json, text/plain, */*",
"X-Requested-With": "XMLHttpRequest",
"X-KL-KIS-Ajax-Request": "Ajax_Request", # Seen in a Yargitay client example
"Referer": f"{self.BASE_URL}/" # Some APIs might check referer
},
timeout=request_timeout,
verify=False # SSL verification disabled as per original user code - use with caution
)
async def search_detailed_decisions(
self,
search_params: YargitayDetailedSearchRequest
) -> YargitayApiSearchResponse:
"""
Performs a detailed search for decisions in Yargitay
using the structured search_params.
"""
# Create the main payload structure with the 'data' key
request_payload = {"data": search_params.model_dump(exclude_none=True, by_alias=True)}
logger.info(f"YargitayOfficialApiClient: Performing detailed search with payload: {request_payload}")
try:
response = await self.http_client.post(self.DETAILED_SEARCH_ENDPOINT, json=request_payload)
response.raise_for_status() # Raise an exception for HTTP 4xx or 5xx status codes
response_json_data = response.json()
logger.debug(f"YargitayOfficialApiClient: Raw API response: {response_json_data}")
# Handle None or empty data response from API
if response_json_data is None:
logger.warning("YargitayOfficialApiClient: API returned None response")
response_json_data = {"data": {"data": [], "recordsTotal": 0, "recordsFiltered": 0}}
elif not isinstance(response_json_data, dict):
logger.warning(f"YargitayOfficialApiClient: API returned unexpected response type: {type(response_json_data)}")
response_json_data = {"data": {"data": [], "recordsTotal": 0, "recordsFiltered": 0}}
elif response_json_data.get("data") is None:
logger.warning("YargitayOfficialApiClient: API response data field is None")
response_json_data["data"] = {"data": [], "recordsTotal": 0, "recordsFiltered": 0}
# Validate and parse the response using Pydantic models
api_response = YargitayApiSearchResponse(**response_json_data)
# Populate the document_url for each decision entry
if api_response.data and api_response.data.data:
for decision_item in api_response.data.data:
decision_item.document_url = f"{self.BASE_URL}{self.DOCUMENT_ENDPOINT}?id={decision_item.id}"
return api_response
except httpx.RequestError as e:
logger.error(f"YargitayOfficialApiClient: HTTP request error during detailed search: {e}")
raise # Re-raise to be handled by the calling MCP tool
except Exception as e: # Catches Pydantic ValidationErrors as well
logger.error(f"YargitayOfficialApiClient: Error processing or validating detailed search response: {e}")
raise
def _convert_html_to_markdown(self, html_from_api_data_field: str) -> Optional[str]:
"""
Takes raw HTML string (from Yargitay API 'data' field for a document),
pre-processes it, and converts it to Markdown using MarkItDown.
Returns only the Markdown string or None if conversion fails.
"""
if not html_from_api_data_field:
return None
# Pre-process HTML: unescape entities and fix common escaped sequences
# Based on user's original fix_html_content
processed_html = html.unescape(html_from_api_data_field)
processed_html = processed_html.replace('\\"', '"')
processed_html = processed_html.replace('\\r\\n', '\n')
processed_html = processed_html.replace('\\n', '\n')
processed_html = processed_html.replace('\\t', '\t')
# MarkItDown often works best with a full HTML document structure.
# The Yargitay /getDokuman response already provides a full <html>...</html> string.
# If it were just a fragment, we might wrap it like:
# html_to_convert = f"<html><head><meta charset=\"UTF-8\"></head><body>{processed_html}</body></html>"
# But since it's already a full HTML string in "data":
html_to_convert = processed_html
markdown_output = None
try:
# Convert HTML string to bytes and create BytesIO stream
html_bytes = html_to_convert.encode('utf-8')
html_stream = io.BytesIO(html_bytes)
# Pass BytesIO stream to MarkItDown to avoid temp file creation
md_converter = MarkItDown()
conversion_result = md_converter.convert(html_stream)
markdown_output = conversion_result.text_content
logger.info("Successfully converted HTML to Markdown.")
except Exception as e:
logger.error(f"Error during MarkItDown HTML to Markdown conversion: {e}")
return markdown_output
async def get_decision_document_as_markdown(self, id: str) -> YargitayDocumentMarkdown:
"""
Retrieves a specific Yargitay decision by its ID and returns its content
as Markdown.
Based on user-provided /getDokuman response structure.
"""
document_api_url = f"{self.DOCUMENT_ENDPOINT}?id={id}"
source_url = f"{self.BASE_URL}{document_api_url}" # The original URL of the document
logger.info(f"YargitayOfficialApiClient: Fetching document for Markdown conversion (ID: {id})")
try:
response = await self.http_client.get(document_api_url)
response.raise_for_status()
# Expecting JSON response with HTML content in the 'data' field
response_json = response.json()
html_content_from_api = response_json.get("data")
if not isinstance(html_content_from_api, str):
logger.error(f"YargitayOfficialApiClient: 'data' field in API response is not a string or not found (ID: {id}).")
raise ValueError("Expected HTML content not found in API response's 'data' field.")
markdown_content = self._convert_html_to_markdown(html_content_from_api)
return YargitayDocumentMarkdown(
id=id,
markdown_content=markdown_content,
source_url=source_url
)
except httpx.RequestError as e:
logger.error(f"YargitayOfficialApiClient: HTTP error fetching document for Markdown (ID: {id}): {e}")
raise
except ValueError as e: # For JSON parsing errors or missing 'data' field
logger.error(f"YargitayOfficialApiClient: Error processing document response for Markdown (ID: {id}): {e}")
raise
except Exception as e: # For other unexpected errors
logger.error(f"YargitayOfficialApiClient: General error fetching/processing document for Markdown (ID: {id}): {e}")
raise
async def close_client_session(self):
"""Closes the HTTPX client session."""
await self.http_client.aclose()
logger.info("YargitayOfficialApiClient: HTTP client session closed.")
@@ -1,103 +0,0 @@
# yargitay_mcp_module/models.py
from pydantic import BaseModel, Field, HttpUrl, ConfigDict
from typing import List, Optional, Dict, Any, Literal
# Yargıtay Chamber/Board Options
YargitayBirimEnum = Literal[
"ALL", # "ALL" for all chambers
# Hukuk (Civil) Chambers
"Hukuk Genel Kurulu",
"1. Hukuk Dairesi", "2. Hukuk Dairesi", "3. Hukuk Dairesi", "4. Hukuk Dairesi",
"5. Hukuk Dairesi", "6. Hukuk Dairesi", "7. Hukuk Dairesi", "8. Hukuk Dairesi",
"9. Hukuk Dairesi", "10. Hukuk Dairesi", "11. Hukuk Dairesi", "12. Hukuk Dairesi",
"13. Hukuk Dairesi", "14. Hukuk Dairesi", "15. Hukuk Dairesi", "16. Hukuk Dairesi",
"17. Hukuk Dairesi", "18. Hukuk Dairesi", "19. Hukuk Dairesi", "20. Hukuk Dairesi",
"21. Hukuk Dairesi", "22. Hukuk Dairesi", "23. Hukuk Dairesi",
"Hukuk Daireleri Başkanlar Kurulu",
# Ceza (Criminal) Chambers
"Ceza Genel Kurulu",
"1. Ceza Dairesi", "2. Ceza Dairesi", "3. Ceza Dairesi", "4. Ceza Dairesi",
"5. Ceza Dairesi", "6. Ceza Dairesi", "7. Ceza Dairesi", "8. Ceza Dairesi",
"9. Ceza Dairesi", "10. Ceza Dairesi", "11. Ceza Dairesi", "12. Ceza Dairesi",
"13. Ceza Dairesi", "14. Ceza Dairesi", "15. Ceza Dairesi", "16. Ceza Dairesi",
"17. Ceza Dairesi", "18. Ceza Dairesi", "19. Ceza Dairesi", "20. Ceza Dairesi",
"21. Ceza Dairesi", "22. Ceza Dairesi", "23. Ceza Dairesi",
"Ceza Daireleri Başkanlar Kurulu",
# General Assembly
"Büyük Genel Kurulu"
]
class YargitayDetailedSearchRequest(BaseModel):
"""
Model for the 'data' object sent in the request payload
to Yargitay's detailed search endpoint (e.g., /aramadetaylist).
Based on the payload provided by the user.
"""
arananKelime: Optional[str] = Field("", description="Turkish keywords (supports +word -word \"phrase\" operators)")
# Department/Board selection - Complete Court of Cassation chamber hierarchy
birimYrgKurulDaire: Optional[str] = Field("ALL", description="Chamber (ALL or specific chamber name)")
esasYil: Optional[str] = Field("", description="Case year (YYYY)")
esasIlkSiraNo: Optional[str] = Field("", description="Start case no")
esasSonSiraNo: Optional[str] = Field("", description="End case no")
kararYil: Optional[str] = Field("", description="Decision year (YYYY)")
kararIlkSiraNo: Optional[str] = Field("", description="Start decision no")
kararSonSiraNo: Optional[str] = Field("", description="End decision no")
baslangicTarihi: Optional[str] = Field("", description="Start date (DD.MM.YYYY)")
bitisTarihi: Optional[str] = Field("", description="End date (DD.MM.YYYY)")
pageSize: int = Field(10, ge=1, le=10, description="Results per page (1-100)")
pageNumber: int = Field(1, ge=1, description="Page number (1-indexed)")
class YargitayApiDecisionEntry(BaseModel):
"""Model for an individual decision entry from the Yargitay API search response."""
id: str # Unique system ID of the decision
daire: Optional[str] = Field(None, description="Chamber")
esasNo: Optional[str] = Field(None, alias="esasNo", description="Case no")
kararNo: Optional[str] = Field(None, alias="kararNo", description="Decision no")
kararTarihi: Optional[str] = Field(None, alias="kararTarihi", description="Date")
# 'index' and 'siraNo' from API response are not critical for MCP tool, so omitted for brevity
# This field will be populated by the client after fetching the search list
document_url: Optional[HttpUrl] = Field(None, description="Document URL")
model_config = ConfigDict(populate_by_name=True) # To allow populating by alias from API response
class YargitayApiResponseInnerData(BaseModel):
"""Model for the inner 'data' object in the Yargitay API search response."""
data: List[YargitayApiDecisionEntry] = Field(default_factory=list)
# draw: Optional[int] = None # Typically used by DataTables, not essential for MCP
recordsTotal: int = Field(default=0) # Total number of records matching the query
recordsFiltered: int = Field(default=0) # Total number of records after filtering (usually same as recordsTotal)
class YargitayApiSearchResponse(BaseModel):
"""Model for the complete search response from the Yargitay API."""
data: Optional[YargitayApiResponseInnerData] = Field(default_factory=lambda: YargitayApiResponseInnerData())
# metadata: Optional[Dict[str, Any]] = None # Optional metadata from API
class YargitayDocumentMarkdown(BaseModel):
"""Model for a Yargitay decision document, containing only Markdown content."""
id: str = Field(..., description="Document ID")
markdown_content: Optional[str] = Field(None, description="Content")
source_url: HttpUrl = Field(..., description="Source URL")
class CleanYargitayDecisionEntry(BaseModel):
"""Clean decision entry without arananKelime field to reduce token usage."""
id: str
daire: Optional[str] = Field(None, description="Chamber")
esasNo: Optional[str] = Field(None, description="Case no")
kararNo: Optional[str] = Field(None, description="Decision no")
kararTarihi: Optional[str] = Field(None, description="Date")
document_url: Optional[HttpUrl] = Field(None, description="Document URL")
class CompactYargitaySearchResult(BaseModel):
"""A more compact search result model for the MCP tool to return."""
decisions: List[CleanYargitayDecisionEntry]
total_records: int
requested_page: int
page_size: int
+7
View File
@@ -0,0 +1,7 @@
# semantic_search/__init__.py
from .embedder import OpenRouterEmbedder, is_openrouter_available
from .vector_store import VectorStore
from .processor import DocumentProcessor
__all__ = ['OpenRouterEmbedder', 'is_openrouter_available', 'VectorStore', 'DocumentProcessor']
+154
View File
@@ -0,0 +1,154 @@
# semantic_search/embedder.py
import logging
import os
from typing import List, Optional
import numpy as np
logger = logging.getLogger(__name__)
def is_openrouter_available() -> bool:
"""Check if OpenRouter API key is available."""
return bool(os.getenv("OPENROUTER_API_KEY"))
class OpenRouterEmbedder:
"""
Embedder using OpenRouter API with Google's Gemini Embedding model.
Requires OPENROUTER_API_KEY environment variable.
"""
def __init__(self):
"""
Initialize OpenRouter Embedder.
Raises:
ValueError: If OPENROUTER_API_KEY is not set
ImportError: If openai package is not installed
"""
api_key = os.getenv("OPENROUTER_API_KEY")
if not api_key:
raise ValueError("OPENROUTER_API_KEY environment variable is not set")
try:
from openai import OpenAI
except ImportError:
raise ImportError("openai package is required. Install with: pip install openai")
self.client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=api_key,
)
self.model = "google/gemini-embedding-001"
self.dimension = 3072
logger.info(f"OpenRouter Embedder initialized with model: {self.model}")
def encode_query(self, query: str, task: str = "search result") -> np.ndarray:
"""
Encode a search query.
Args:
query: The search query text
task: Task type for prompt template
Returns:
Numpy array of embeddings (3072 dimensions)
"""
# Apply query prompt template
text = f"task: {task} | query: {query}"
try:
response = self.client.embeddings.create(
model=self.model,
input=text,
encoding_format="float",
extra_headers={
"HTTP-Referer": "https://yargimcp.com",
"X-Title": "Yargi MCP Server",
}
)
embedding = np.array(response.data[0].embedding, dtype=np.float32)
# L2 normalize for cosine similarity
norm = np.linalg.norm(embedding)
if norm > 0:
embedding = embedding / norm
logger.debug(f"Encoded query: {query[:50]}... -> shape: {embedding.shape}")
return embedding
except Exception as e:
logger.error(f"Failed to encode query: {e}")
raise
def encode_documents(self, documents: List[str], titles: Optional[List[str]] = None) -> np.ndarray:
"""
Encode multiple documents with batch API call.
Args:
documents: List of document texts
titles: Optional list of document titles
Returns:
Numpy array of embeddings (N x 3072 dimensions)
"""
if not documents:
return np.array([])
# Apply document prompt template
texts = []
for i, doc in enumerate(documents):
title = titles[i] if titles and i < len(titles) else "none"
text = f"title: {title} | text: {doc}"
texts.append(text)
try:
response = self.client.embeddings.create(
model=self.model,
input=texts,
encoding_format="float",
extra_headers={
"HTTP-Referer": "https://yargimcp.com",
"X-Title": "Yargi MCP Server",
}
)
# Extract embeddings in order
embeddings = np.array(
[d.embedding for d in sorted(response.data, key=lambda x: x.index)],
dtype=np.float32
)
# L2 normalize each embedding for cosine similarity
norms = np.linalg.norm(embeddings, axis=1, keepdims=True)
embeddings = embeddings / (norms + 1e-8)
logger.info(f"Encoded {len(documents)} documents -> shape: {embeddings.shape}")
return embeddings
except Exception as e:
logger.error(f"Failed to encode documents: {e}")
raise
def compute_similarity(self, query_embedding: np.ndarray, document_embeddings: np.ndarray) -> np.ndarray:
"""
Compute cosine similarity between query and documents.
Args:
query_embedding: Query embedding (3072,)
document_embeddings: Document embeddings (N x 3072)
Returns:
Similarity scores (N,)
"""
# Ensure query is 2D for matrix multiplication
if len(query_embedding.shape) == 1:
query_embedding = query_embedding.reshape(1, -1)
# Compute cosine similarity (embeddings are already normalized)
similarities = np.dot(document_embeddings, query_embedding.T).squeeze()
return similarities
+305
View File
@@ -0,0 +1,305 @@
# semantic_search/processor.py
import logging
import re
from typing import List, Dict, Any, Optional
from dataclasses import dataclass
import hashlib
logger = logging.getLogger(__name__)
@dataclass
class DocumentChunk:
"""Represents a chunk of a document."""
chunk_id: str
document_id: str
text: str
metadata: Dict[str, Any]
chunk_index: int
total_chunks: int
class DocumentProcessor:
"""
Processes legal documents for semantic search.
Handles chunking, cleaning, and metadata extraction.
"""
def __init__(self,
chunk_size: int = 1000,
chunk_overlap: int = 200,
min_chunk_size: int = 100):
"""
Initialize document processor.
Args:
chunk_size: Target size for each chunk in characters
chunk_overlap: Number of overlapping characters between chunks
min_chunk_size: Minimum chunk size to keep
"""
self.chunk_size = chunk_size
self.chunk_overlap = chunk_overlap
self.min_chunk_size = min_chunk_size
logger.info(f"Initialized DocumentProcessor (chunk_size={chunk_size}, overlap={chunk_overlap})")
def process_document(self,
document_id: str,
text: str,
metadata: Optional[Dict[str, Any]] = None) -> List[DocumentChunk]:
"""
Process a single document into chunks.
Args:
document_id: Unique document identifier
text: Document text content
metadata: Optional document metadata
Returns:
List of document chunks
"""
if not text or len(text.strip()) < self.min_chunk_size:
logger.warning(f"Document {document_id} too short to process")
return []
# Clean text
cleaned_text = self._clean_text(text)
# Extract metadata from text if not provided
if metadata is None:
metadata = {}
# Add extracted metadata
extracted_metadata = self._extract_metadata(cleaned_text)
metadata.update(extracted_metadata)
# Create chunks
chunks = self._create_chunks(cleaned_text)
# Create DocumentChunk objects
document_chunks = []
for i, chunk_text in enumerate(chunks):
chunk_id = self._generate_chunk_id(document_id, i)
chunk = DocumentChunk(
chunk_id=chunk_id,
document_id=document_id,
text=chunk_text,
metadata={
**metadata,
'chunk_index': i,
'total_chunks': len(chunks)
},
chunk_index=i,
total_chunks=len(chunks)
)
document_chunks.append(chunk)
logger.info(f"Processed document {document_id} into {len(chunks)} chunks")
return document_chunks
def _clean_text(self, text: str) -> str:
"""
Clean and normalize text for processing.
Args:
text: Raw text
Returns:
Cleaned text
"""
# Remove excessive whitespace
text = re.sub(r'\s+', ' ', text)
# Remove special characters but keep Turkish characters
# Keep: letters, numbers, spaces, and common punctuation
text = re.sub(r'[^\w\s\.\,\;\:\!\?\-\(\)\"\'ÇĞIİÖŞÜçğıiöşü]', ' ', text)
# Remove multiple spaces
text = re.sub(r' +', ' ', text)
# Trim
text = text.strip()
return text
def _extract_metadata(self, text: str) -> Dict[str, Any]:
"""
Extract metadata from legal document text.
Args:
text: Document text
Returns:
Extracted metadata
"""
metadata = {}
# Extract case numbers (Esas/Karar)
esas_pattern = r'E(?:sas)?[\s\.\:]*(\d{4})[\/\-](\d+)'
karar_pattern = r'K(?:arar)?[\s\.\:]*(\d{4})[\/\-](\d+)'
esas_match = re.search(esas_pattern, text[:500]) # Look in first 500 chars
if esas_match:
metadata['esas_no'] = f"E.{esas_match.group(1)}/{esas_match.group(2)}"
karar_match = re.search(karar_pattern, text[:500])
if karar_match:
metadata['karar_no'] = f"K.{karar_match.group(1)}/{karar_match.group(2)}"
# Extract dates (DD.MM.YYYY or DD/MM/YYYY format)
date_pattern = r'(\d{1,2})[\.\/](\d{1,2})[\.\/](\d{4})'
dates = re.findall(date_pattern, text[:1000]) # Look in first 1000 chars
if dates:
# Take the first date as decision date
day, month, year = dates[0]
metadata['karar_tarihi'] = f"{year}-{month.zfill(2)}-{day.zfill(2)}"
# Extract court/chamber name
chamber_patterns = [
r'(\d+)\.\s*Hukuk\s+Dairesi',
r'(\d+)\.\s*Ceza\s+Dairesi',
r'Hukuk\s+Genel\s+Kurulu',
r'Ceza\s+Genel\s+Kurulu',
r'(\d+)\.\s*Daire'
]
for pattern in chamber_patterns:
match = re.search(pattern, text[:500], re.IGNORECASE)
if match:
metadata['chamber'] = match.group(0)
break
return metadata
def _create_chunks(self, text: str) -> List[str]:
"""
Create overlapping chunks from text.
Args:
text: Cleaned document text
Returns:
List of text chunks
"""
chunks = []
# Split by sentences for better semantic coherence
sentences = self._split_sentences(text)
current_chunk = []
current_size = 0
for sentence in sentences:
sentence_size = len(sentence)
# If adding this sentence exceeds chunk size
if current_size + sentence_size > self.chunk_size and current_chunk:
# Save current chunk
chunk_text = ' '.join(current_chunk)
chunks.append(chunk_text)
# Create overlap for next chunk
overlap_size = 0
overlap_sentences = []
# Add sentences from the end until we reach overlap size
for sent in reversed(current_chunk):
overlap_size += len(sent)
overlap_sentences.insert(0, sent)
if overlap_size >= self.chunk_overlap:
break
# Start new chunk with overlap
current_chunk = overlap_sentences
current_size = sum(len(s) for s in current_chunk)
# Add sentence to current chunk
current_chunk.append(sentence)
current_size += sentence_size
# Add final chunk if not empty
if current_chunk:
chunk_text = ' '.join(current_chunk)
if len(chunk_text) >= self.min_chunk_size:
chunks.append(chunk_text)
return chunks
def _split_sentences(self, text: str) -> List[str]:
"""
Split text into sentences.
Args:
text: Text to split
Returns:
List of sentences
"""
# Simple sentence splitting for Turkish text
# Split on period, question mark, exclamation, but not on abbreviations
# Common Turkish abbreviations to preserve
abbreviations = ['Dr', 'Prof', 'Av', 'Md', 'Yrd', 'Doç', 'No', 'S', 'vs', 'vb', 'bkz']
# Replace abbreviations temporarily
temp_text = text
replacements = {}
for i, abbr in enumerate(abbreviations):
placeholder = f"__ABBR{i}__"
temp_text = temp_text.replace(f"{abbr}.", placeholder)
replacements[placeholder] = f"{abbr}."
# Split sentences
sentence_endings = re.compile(r'[.!?]+')
sentences = sentence_endings.split(temp_text)
# Restore abbreviations and clean
cleaned_sentences = []
for sentence in sentences:
# Restore abbreviations
for placeholder, original in replacements.items():
sentence = sentence.replace(placeholder, original)
# Clean and add if not empty
sentence = sentence.strip()
if sentence and len(sentence) > 10: # Minimum sentence length
cleaned_sentences.append(sentence)
return cleaned_sentences
def _generate_chunk_id(self, document_id: str, chunk_index: int) -> str:
"""
Generate unique chunk ID.
Args:
document_id: Parent document ID
chunk_index: Index of chunk in document
Returns:
Unique chunk ID
"""
chunk_string = f"{document_id}_chunk_{chunk_index}"
chunk_hash = hashlib.md5(chunk_string.encode()).hexdigest()[:8]
return f"{document_id}_c{chunk_index}_{chunk_hash}"
def combine_chunks(self, chunks: List[DocumentChunk]) -> str:
"""
Combine chunks back into full document text.
Args:
chunks: List of document chunks
Returns:
Combined text
"""
if not chunks:
return ""
# Sort by chunk index
sorted_chunks = sorted(chunks, key=lambda x: x.chunk_index)
# For overlapping chunks, we need to be careful about duplication
# Simple approach: just concatenate with space
combined = " ".join([chunk.text for chunk in sorted_chunks])
return combined
+235
View File
@@ -0,0 +1,235 @@
# semantic_search/vector_store.py
import logging
import numpy as np
from typing import List, Dict, Any, Tuple, Optional
from dataclasses import dataclass
import json
logger = logging.getLogger(__name__)
@dataclass
class Document:
"""Represents a document with its embedding and metadata."""
id: str
text: str
embedding: np.ndarray
metadata: Dict[str, Any]
def to_dict(self) -> Dict[str, Any]:
"""Convert to dictionary (excluding embedding for serialization)."""
return {
'id': self.id,
'text': self.text,
'metadata': self.metadata
}
class VectorStore:
"""
In-memory vector storage with similarity search capabilities.
Future versions can use Faiss, ChromaDB, or other vector databases.
"""
def __init__(self, dimension: int = 768):
"""
Initialize vector store.
Args:
dimension: Embedding dimension size
"""
self.dimension = dimension
self.documents: List[Document] = []
self.embeddings: Optional[np.ndarray] = None
self.index_built = False
logger.info(f"Initialized VectorStore with dimension: {dimension}")
def add_documents(self,
ids: List[str],
texts: List[str],
embeddings: np.ndarray,
metadata: Optional[List[Dict[str, Any]]] = None) -> int:
"""
Add documents to the vector store.
Args:
ids: Document IDs
texts: Document texts
embeddings: Document embeddings (N x dimension)
metadata: Optional metadata for each document
Returns:
Number of documents added
"""
if len(ids) != len(texts) or len(ids) != embeddings.shape[0]:
raise ValueError("Mismatched lengths for ids, texts, and embeddings")
if metadata and len(metadata) != len(ids):
raise ValueError("Metadata length doesn't match document count")
# Add documents
for i in range(len(ids)):
doc = Document(
id=ids[i],
text=texts[i],
embedding=embeddings[i],
metadata=metadata[i] if metadata else {}
)
self.documents.append(doc)
# Rebuild index
self._build_index()
logger.info(f"Added {len(ids)} documents to vector store. Total: {len(self.documents)}")
return len(ids)
def _build_index(self):
"""Build or rebuild the embedding index."""
if not self.documents:
self.embeddings = None
self.index_built = False
return
# Stack all embeddings into a single array
self.embeddings = np.vstack([doc.embedding for doc in self.documents])
self.index_built = True
logger.debug(f"Built index with shape: {self.embeddings.shape}")
def search(self,
query_embedding: np.ndarray,
top_k: int = 10,
threshold: Optional[float] = None) -> List[Tuple[Document, float]]:
"""
Search for similar documents using cosine similarity.
Args:
query_embedding: Query embedding vector
top_k: Number of results to return
threshold: Optional similarity threshold (0-1)
Returns:
List of (Document, similarity_score) tuples
"""
if not self.index_built or self.embeddings is None:
logger.warning("No documents in vector store")
return []
# Ensure query is 2D
if len(query_embedding.shape) == 1:
query_embedding = query_embedding.reshape(1, -1)
# Compute cosine similarities (assuming normalized embeddings)
similarities = np.dot(self.embeddings, query_embedding.T).squeeze()
# Apply threshold if specified
if threshold is not None:
valid_indices = np.where(similarities >= threshold)[0]
if len(valid_indices) == 0:
logger.info(f"No documents above threshold {threshold}")
return []
similarities = similarities[valid_indices]
valid_docs = [self.documents[i] for i in valid_indices]
else:
valid_docs = self.documents
# Get top-k indices
top_k = min(top_k, len(valid_docs))
if top_k == 0:
return []
# Use argpartition for efficiency with large arrays
if len(similarities) > top_k:
top_indices = np.argpartition(similarities, -top_k)[-top_k:]
top_indices = top_indices[np.argsort(similarities[top_indices])[::-1]]
else:
top_indices = np.argsort(similarities)[::-1]
# Create results
results = []
for idx in top_indices:
doc = valid_docs[idx] if threshold else self.documents[idx]
score = float(similarities[idx])
results.append((doc, score))
logger.info(f"Search returned {len(results)} results (top_k={top_k})")
return results
def hybrid_search(self,
query_embedding: np.ndarray,
keyword_scores: Dict[str, float],
top_k: int = 10,
alpha: float = 0.5) -> List[Tuple[Document, float]]:
"""
Hybrid search combining vector similarity and keyword scores.
Args:
query_embedding: Query embedding vector
keyword_scores: Document ID to keyword relevance score mapping
top_k: Number of results to return
alpha: Weight for vector similarity (1-alpha for keyword score)
Returns:
List of (Document, combined_score) tuples
"""
if not self.index_built:
logger.warning("No documents in vector store")
return []
# Get vector similarities
vector_results = self.search(query_embedding, top_k=len(self.documents))
# Combine scores
combined_scores = []
for doc, vector_score in vector_results:
keyword_score = keyword_scores.get(doc.id, 0.0)
# Normalize keyword score to 0-1 range if needed
if keyword_score > 1.0:
keyword_score = keyword_score / max(keyword_scores.values())
combined_score = alpha * vector_score + (1 - alpha) * keyword_score
combined_scores.append((doc, combined_score))
# Sort by combined score and return top-k
combined_scores.sort(key=lambda x: x[1], reverse=True)
results = combined_scores[:top_k]
logger.info(f"Hybrid search returned {len(results)} results")
return results
def clear(self):
"""Clear all documents from the store."""
self.documents = []
self.embeddings = None
self.index_built = False
logger.info("Cleared vector store")
def size(self) -> int:
"""Get number of documents in store."""
return len(self.documents)
def get_by_id(self, doc_id: str) -> Optional[Document]:
"""Get document by ID."""
for doc in self.documents:
if doc.id == doc_id:
return doc
return None
def get_stats(self) -> Dict[str, Any]:
"""Get statistics about the vector store."""
stats = {
'num_documents': len(self.documents),
'dimension': self.dimension,
'index_built': self.index_built,
'memory_usage_mb': 0
}
if self.embeddings is not None:
# Estimate memory usage
memory_bytes = self.embeddings.nbytes
for doc in self.documents:
memory_bytes += len(doc.text.encode('utf-8'))
memory_bytes += len(json.dumps(doc.metadata).encode('utf-8'))
stats['memory_usage_mb'] = memory_bytes / (1024 * 1024)
return stats
Generated
+1315 -1193
View File
File diff suppressed because it is too large Load Diff