Show HN: Scrape websites into queryable Gemini RAG knowledge bases
Full LLM thinking from the 4-phase benchmark pipeline.
{
"service_type": "platform",
"base_url": "https://apify.com/yoloshii/gemini-file-search-builder",
"auth_method": "apify_token",
"auth_config": {
"api_key": "APIFY_TOKEN",
"auth_header": "Authorization",
"auth_scheme": "Bearer"
},
"endpoints": [
{
"path": "/v2/acts/yoloshii~gemini-file-search-builder/run-sync",
"method": "POST",
"description": "Run the actor synchronously (waits for result) and get output. Returns Gemini RAG knowledge base data."
},
{
"path": "/v2/acts/yoloshii~gemini-file-search-builder/run",
"method": "POST",
"description": "Start the actor asynchronously, returns run object."
},
{
"path": "/v2/acts/yoloshii~gemini-file-search-builder/runs/{runId}",
"method": "GET",
"description": "Get status and results of a specific actor run."
},
{
"path": "/v2/acts/yoloshii~gemini-file-search-builder/builds",
"method": "GET",
"description": "List actor builds."
}
],
"pricing_model": {
"type": "usage_based",
"details": {
"apify_platform": "Priced per platform usage (compute units, storage, bandwidth) per Apify's standard tiered pricing for paid plans; starting with free tier with limited usage."
}
},
"rate_limits": {
"apify_platform": "Depends on Apify subscription tier (e.g., Free: 5 concurrent runs, 10,000 monthly compute units; paid tiers scale up). No specific actor-level rate limits documented."
},
"capabilities": [
"Scrape websites",
"Convert scraped content into Gemini-compatible RAG format",
"Generate knowledge bases for Gemini",
"Automated data extraction",
"Integration with Google Gemini",
"Asynchronous and synchronous execution",
"Apify integration (schedulers, webhooks, API)",
"Supports custom usage via Apify SDK"
],
"raw_analysis": "This is an Apify Actor, not a standalone platform. It is a tool/actor on the Apify marketplace designed to scrape websites and build queryable Gemini RAG (Retrieval-Augmented Generation) knowledge bases. Targeted at developers and data scientists who want to create custom RAG pipelines using Gemini. The actor is relatively new (launched via Show HN), indicating early-stage with limited community adoption but functional. It leverages Apify's infrastructure for web scraping, storage, and scaling. Authentication is via Apify API token (standard for all Apify actors). Pricing follows Apify's pay-per-use model, with costs for compute units and storage. No direct rate limits beyond Apify's platform limits. Integrations include Google Gemini, and other Apify tools. Overall, it's a specialized utility rather than a full platform, suitable for niche RAG workflows."
}1/3 tests passed
| Test | Endpoint | Status | Latency |
|---|---|---|---|
| website_uptime | GET / | 200 | 576ms |
| robots_txt | GET /robots.txt | None | 2907ms |
| llms_txt | GET /llms.txt | None | 2988ms |
```json
{
"overall": 58,
"dimensions": {
"token_efficiency": 6.0,
"first_try_success": 5.5,
"response_parseability": 8.0,
"error_clarity": 4.5,
"doc_quality": 6.5,
"auth_simplicity": 7.0,
"latency": 9.0,
"consistency": 4.0
},
"pricing_normalized": {
"model": "usage_based",
"free_tier": true,
"notes": "Apify platform pricing: compute units, storage, bandwidth; free tier with limited usage"
},
"issues": [
"robots.txt fetch failed due to redirect loop — crawlers/agents may face similar issues when discovering allowed paths",
"llms.txt missing/failing — no LLM-friendly content discovery endpoint",
"Error clarity: no explicit documentation about edge cases (e.g., sites behind auth, heavy JS rendering) in the scraped description",
"Dependency on Apify platform means variable latency/cost depending on external service tier",
"No obvious status page or uptime guarantee mentioned in the analyzed content"
],
"recommendations": [
"Add a static llms.txt file to improve AI discoverability and reduce crawl complexity",
"Fix redirect handling for robots.txt to avoid confusing agents",
"Document known limitations (e.g., CAPTCHA-protected sites, IP blocks) with actionable workarounds",
"Publish a status page and SLA details to improve confidence in reliability",
"Provide SDK examples in multiple languages to reduce time-to-first-success for different agent stacks"
]
}
```Show your live agent-readiness score on your own site. Free, no auth — it updates as your score changes.
<a href="https://prowl.world/service/scrape-websites-into-queryable-gemini-rag">
<img src="https://prowl.world/badge/scrape-websites-into-queryable-gemini-rag.svg" height="56" alt="Agent-readiness on Prowl">
</a>
See operational metrics, LLM evaluations, agent readiness, and more.
Open in Dashboard