diff --git a/CHANGELOG.md b/CHANGELOG.md index 8fbbe7c7..3bf1e4b8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -65,3 +65,35 @@ Updated every Monday. --- *Next update: 2026-03-30* + +## [v0.7.0] — 2026-03-31 + +### 🚀 大规模扩充:388 → 560+ 精选 Skills(+173) + +来源:ClawHub Archive (openclaw/skills) + GitHub 精选 + +#### 新增亮点 + +**steipete (OpenClaw 创始人) 新 Skills (13个)** +- `bird` — Bluesky 社交协议 CLI +- `coding-agent` — AI 编码 agent 最佳实践 +- `create-cli` — 创建命令行工具 +- `domain-dns-ops` — 域名 DNS 管理 +- `frontend-design` — 前端设计指南 +- `instruments-profiling` — Xcode Instruments 性能分析 +- `native-app-performance` — 原生 App 性能优化 +- `swift-concurrency-expert` — Swift 并发专家 +- `swiftui-liquid-glass` — SwiftUI Liquid Glass 效果 +- `swiftui-performance-audit` — SwiftUI 性能审计 +- `swiftui-view-refactor` — SwiftUI 视图重构 +- `video-transcript-downloader` — 视频字幕下载 +- `wacli` — WhatsApp CLI + +**alirezarezvani 企业级 Skills (159个)** +覆盖:C-Suite 顾问(CEO/CFO/CTO/CMO/CISO等)、高级工程师(Frontend/Backend/ML/DevOps/Security)、产品管理、项目管理、UI/UX 设计、内容创作、SEO/SEM、社交媒体、数据分析、Terraform/K8s/Docker、Stripe 集成等 + +**其他新增** +- `para-second-brain` — PARA 方法论第二大脑系统 + +--- + diff --git a/README.zh-CN.md b/README.zh-CN.md index 4d9709c8..0beded8b 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -5,7 +5,7 @@ Powered by MyClaw.ai -339+ Skills +560+ Skills Weekly Updates **语言:** @@ -34,7 +34,7 @@ git clone https://github.com/LeoYeAI/openclaw-master-skills.git cp -r openclaw-master-skills/skills/ ~/.openclaw/workspace/skills/ ``` -## 📦 技能索引 (339 skills) +## 📦 技能索引 (560 skills) ### 🤖 AI & LLM Tools (34) diff --git a/RELEASES.md b/RELEASES.md index 899e04cd..e306c9a0 100644 --- a/RELEASES.md +++ b/RELEASES.md @@ -413,3 +413,33 @@ - `vercel-react-best-practices` — React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, r - `web-design-guidelines` — Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit - `supabase-postgres-best-practices` — Postgres performance optimization and best practices from Supabase. Use this skill when writing, reviewing, or optimizin + +--- + +## v0.7.0 — 2026-03-31 + +### 🚀 Weekly Update — 560+ Skills (+173) + +Sources: ClawHub Archive (openclaw/skills repo) + GitHub curation + +#### New from steipete (OpenClaw creator) — 13 skills +- `bird` — Bluesky social protocol CLI +- `coding-agent` — AI coding agent best practices +- `create-cli` — CLI tool creation guide +- `domain-dns-ops` — Domain DNS operations +- `frontend-design` — Frontend design guidelines +- `instruments-profiling` — Xcode Instruments profiling +- `native-app-performance` — Native app performance optimization +- `swift-concurrency-expert` — Swift concurrency patterns +- `swiftui-liquid-glass` — SwiftUI Liquid Glass effects +- `swiftui-performance-audit` — SwiftUI performance auditing +- `swiftui-view-refactor` — SwiftUI view refactoring +- `video-transcript-downloader` — Video transcript download +- `wacli` — WhatsApp CLI + +#### New from alirezarezvani — 159 enterprise-grade skills +Categories: C-Suite advisors (CEO/CFO/CTO/CMO/CISO), Senior engineers (Frontend/Backend/ML/DevOps/Security), Product management, Project management, UI/UX design, Content creation, SEO/SEM, Social media, Data analytics, Terraform/K8s/Docker, Stripe integration, and more. + +#### Other additions +- `para-second-brain` — PARA method second brain system + diff --git a/SKILL.md b/SKILL.md index 503f90b4..e2e99cbd 100644 --- a/SKILL.md +++ b/SKILL.md @@ -1,13 +1,13 @@ --- name: openclaw-master-skills -description: "A curated collection of 339+ best OpenClaw skills — AI tools, productivity, marketing, frontend, mobile, backend, DevOps and more. Weekly updated by MyClaw.ai — Powered by MyClaw.ai" +description: "A curated collection of 560+ best OpenClaw skills — AI tools, productivity, marketing, frontend, mobile, backend, DevOps and more. Weekly updated by MyClaw.ai — Powered by MyClaw.ai" metadata: openclaw: {} --- # OpenClaw Master Skills -A curated, weekly-updated collection of **339+ best skills** for OpenClaw agents. +A curated, weekly-updated collection of **560+ best skills** for OpenClaw agents. ## Categories diff --git a/pending/audit-website/README.md b/pending/audit-website/README.md deleted file mode 100644 index 9cdeeaff..00000000 --- a/pending/audit-website/README.md +++ /dev/null @@ -1,20 +0,0 @@ -![squirrelscan](https://mintcdn.com/squirrelscan/CCMTmLbI4xfnpJbQ/logo/light.svg?fit=max&auto=format&n=CCMTmLbI4xfnpJbQ&q=85&s=1303484a4ea3c154c29dd5f6245e55cd) - -# squirrelscan Skills - -**CLI Website Audits for Humans, Agents & LLMs** - -## What is squirrelscan? - -[squirrelscan](https://squirrelscan.com) is a comprehensive website audit tool designed for developers, SEO professionals, and AI coding assistants. Built specifically to integrate seamlessly into modern development workflows and AI-assisted coding environments. - -**Features:** -- 200+ audit rules across SEO, performance, accessibility, content, and security -- Leaked secrets detection (96 patterns: OpenAI, Anthropic, AWS, Stripe, and more) -- Multiple output formats: console, text, json, markdown, llm, html -- Diff reports for regressions between audits -- Designed for both human developers and AI agents -- Optimized for CI/CD pipelines and automation -- LLM-native output for AI-assisted debugging and optimization - -Whether you're debugging SEO issues, validating site health, or enabling your AI assistant to autonomously fix website problems, squirrelscan fits into your workflow. diff --git a/pending/audit-website/SKILL.md b/pending/audit-website/SKILL.md deleted file mode 100644 index 7e796096..00000000 --- a/pending/audit-website/SKILL.md +++ /dev/null @@ -1,470 +0,0 @@ ---- -name: audit-website -description: Audit websites for SEO, performance, security, technical, content, and 15 other issue cateories with 230+ rules using the squirrelscan CLI. Returns LLM-optimized reports with health scores, broken links, meta tag analysis, and actionable recommendations. Use to discover and asses website or webapp issues and health. -license: See LICENSE file in repository root -compatibility: Requires squirrel CLI installed and accessible in PATH -metadata: - author: squirrelscan - version: "1.22" -allowed-tools: Bash(squirrel:*) Read Edit Grep Glob ---- - -# Website Audit Skill - -Audit websites for SEO, technical, content, performance and security issues using the squirrelscan cli. - -squirrelscan provides a cli tool squirrel - available for macos, windows and linux. It carries out extensive website auditing -by emulating a browser, search crawler, and analyzing the website's structure and content against over 230+ rules. - -It will provide you a list of issues as well as suggestions on how to fix them. - -## Links - -* squirrelscan website is at [https://squirrelscan.com](https://squirrelscan.com) -* documentation (including rule references) are at [docs.squirrelscan.com](https://docs.squirrelscan.com) - -You can look up the docs for any rule with this template: - -https://docs.squirrelscan.com/rules/{rule_category}/{rule_id} - -example: - -https://docs.squirrelscan.com/rules/links/external-links - -## What This Skill Does - -This skill enables AI agents to audit websites for over 230 rules in 21 categories, including: - -- **SEO issues**: Meta tags, titles, descriptions, canonical URLs, Open Graph tags -- **Technical problems**: Broken links, redirect chains, page speed, mobile-friendliness -- **Performance**: Page load time, resource usage, caching -- **Content quality**: Heading structure, image alt text, content analysis -- **Security**: Leaked secrets, HTTPS usage, security headers, mixed content -- **Accessibility**: Alt text, color contrast, keyboard navigation -- **Usability**: Form validation, error handling, user flow -- **Links**: Checks for broken internal and external links -- **E-E-A-T**: Expertise, Experience, Authority, Trustworthiness -- **User Experience**: User flow, error handling, form validation -- **Mobile**: Checks for mobile-friendliness, responsive design, touch-friendly elements -- **Crawlability**: Checks for crawlability, robots.txt, sitemap.xml and more -- **Schema**: Schema.org markup, structured data, rich snippets -- **Legal**: Compliance with legal requirements, privacy policies, terms of service -- **Social**: Open graph, twitter cards and validating schemas, snippets etc. -- **Url Structure**: Length, hyphens, keywords -- **Keywords**: Keyword stuffing -- **Content**: Content structure, headings -- **Images**: Alt text, color contrast, image size, image format -- **Local SEO**: NAP consistency, geo metadata -- **Video**: VideoObject schema, accessibility - -and more - -The audit crawls the website, analyzes each page against audit rules, and returns a comprehensive report with: -- Overall health score (0-100) -- Category breakdowns (core SEO, technical SEO, content, security) -- Specific issues with affected URLs -- Broken link detection -- Actionable recommendations -- Rules have levels of error, warning and notice and also have a rank between 1 and 10 - -## When to Use - -Use this skill when you need to: - -- Analyze a website's health -- Debug technical SEO issues -- Fix all of the issues mentioned above -- Check for broken links -- Validate meta tags and structured data -- Generate site audit reports -- Compare site health before/after changes -- Improve website performance, accessibility, SEO, security and more. - -You should re-audit as often as possible to ensure your website remains healthy and performs well. - -## Prerequisites - -This skill requires the squirrel CLI installed and in PATH. - -**Install:** [squirrelscan.com/download](https://squirrelscan.com/download) - -**Verify:** -```bash -squirrel --version -``` - -## Setup - -Run `squirrel init` to create a `squirrel.toml` config in the current directory. If none exists, create one and specify a project name: - -```bash -squirrel init -n my-project -# overwrite existing config -squirrel init -n my-project --force -``` - -## Usage - -### Intro - -There are three processes that you can run and they're all cached in the local project database: - -- crawl - subcommand to run a crawl or refresh, continue a crawl -- analyze - subcommand to analyze the crawl results -- report - subcommand to generate a report in desired format (llm, text, console, html etc.) - -the 'audit' command is a wrapper around these three processes and runs them sequentially: - -```bash -squirrel audit https://example.com --format llm -``` - -YOU SHOULD always prefer format option llm - it was made for you and provides an exhaustive and compact output format. - -FIRST SCAN should be a surface scan, which is a quick and shallow scan of the website to gather basic information about the website, such as its structure, content, and technology stack. This scan can be done quickly and without impacting the website's performance. - -SECOND SCAN should be a deep scan, which is a thorough and detailed scan of the website to gather more information about the website, such as its security, performance, and accessibility. This scan can take longer and may impact the website's performance. - -If the user doesn't provide a website to audit, ask which URL they'd like audited. - -You should PREFER to audit live websites - only there do we get a TRUE representation of the website and performance or rendering issuers. - -If you have both local and live websites to audit, prompt the user to choose which one to audit and SUGGEST they choose live. - -You can apply fixes from an audit on the live site against the local code. - -When planning scope tasks so they can run concurrently as sub-agents to speed up fixes. - -When implementing fixes take advantage of subagents to speed up implementation of fixes. - -After applying fixes, verify the code still builds and passes any existing checks in the project. - -### Basic Workflow - -The audit process is two steps: - -1. **Run the audit** (saves to database, shows console output) -2. **Export report** in desired format - -```bash -# Step 1: Run audit (default: console output) -squirrel audit https://example.com - -# Step 2: Export as LLM format -squirrel report --format llm -``` - -### Regression Diffs - -When you need to detect regressions between audits, use diff mode: - -```bash -# Compare current report against a baseline audit ID -squirrel report --diff --format llm - -# Compare latest domain report against a baseline domain -squirrel report --regression-since example.com --format llm -``` - -Diff mode supports `console`, `text`, `json`, `llm`, and `markdown`. `html` and `xml` are not supported. - -### Running Audits - -When running an audit: - -1. **Present the report** - show the user the audit results and score -2. **Propose fixes** - list the issues you can fix and ask the user to confirm before making changes -3. **Parallelize approved fixes** - use subagents for bulk content edits (alt text, headings, descriptions) -4. **Iterate** - fix batch → re-audit → present results → propose next batch -5. **Pause for judgment** - broken links, structural changes, and anything ambiguous should be flagged for user review -6. **Show before/after** - present score comparison after each fix batch - -- **Iteration Loop**: After fixing a batch of issues, re-audit and continue fixing until: - - Score reaches target (typically 85+), OR - - Only issues requiring human judgment remain (e.g., "should this link be removed?") - -- **Treat all fixes equally**: Code changes and content changes are equally important. - -- **Parallelize content fixes**: For issues affecting multiple files: - - Spawn subagents to fix in parallel - - Example: 7 files need alt text → spawn 1-2 agents to fix all - - Example: 30 files have heading issues → spawn agents to batch edit - -- **Completion criteria**: - - ✅ All errors fixed - - ✅ All warnings fixed (or documented as requiring human review) - - ✅ Re-audit confirms improvements - - ✅ Before/after comparison shown to user - -After fixes are applied, ask the user if they'd like to review the changes. - -### Score Targets - -| Starting Score | Target Score | Expected Work | -|----------------|--------------|---------------| -| < 50 (Grade F) | 75+ (Grade C) | Major fixes | -| 50-70 (Grade D) | 85+ (Grade B) | Moderate fixes | -| 70-85 (Grade C) | 90+ (Grade A) | Polish | -| > 85 (Grade B+) | 95+ | Fine-tuning | - -A site is only considered COMPLETE and FIXED when scores are above 95 (Grade A) with coverage set to FULL (--coverage full). - -### Issue Categories - -| Category | Fix Approach | Parallelizable | -|----------|--------------|----------------| -| Meta tags/titles | Edit page components or metadata | No | -| Structured data | Add JSON-LD to page templates | No | -| Missing H1/headings | Edit page components + content files | Yes (content) | -| Image alt text | Edit content files | Yes | -| Heading hierarchy | Edit content files | Yes | -| Short descriptions | Edit content frontmatter | Yes | -| HTTP→HTTPS links | Find and replace in content | Yes | -| Broken links | Manual review (flag for user) | No | - -**For parallelizable fixes**: Spawn subagents with specific file assignments. - -### Content File Fixes - -Many issues require editing content files. These are equally important as code fixes: - -- **Image alt text**: Add descriptive alt text to images -- **Heading hierarchy**: Fix skipped heading levels -- **Meta descriptions**: Extend short descriptions in frontmatter -- **HTTP links**: Update insecure links to HTTPS - -### Parallelizing Fixes with Subagents - -When the user approves a batch of fixes, you can use subagents to apply them in parallel: - -- **Ask the user first** — always confirm which fixes to apply before spawning subagents -- Group 3-5 files per subagent for the same fix type -- Only parallelize independent files (no shared components or config) -- Spawn multiple subagents in a single message for concurrent execution - -### Advanced Options - -Audit more pages: - -```bash -squirrel audit https://example.com --max-pages 200 -``` - -Force fresh crawl (ignore cache): - -```bash -squirrel audit https://example.com --refresh -``` - -Resume interrupted crawl: - -```bash -squirrel audit https://example.com --resume -``` - -Verbose output for debugging: - -```bash -squirrel audit https://example.com --verbose -``` - -## Common Options - -### Audit Command Options - -| Option | Alias | Description | Default | -|--------|-------|-------------|---------| -| `--format ` | `-f ` | Output format: console, text, json, html, markdown, llm | console | -| `--coverage ` | `-C ` | Coverage mode: quick, surface, full | surface | -| `--max-pages ` | `-m ` | Maximum pages to crawl (max 5000) | varies by coverage | -| `--output ` | `-o ` | Output file path | - | -| `--refresh` | `-r` | Ignore cache, fetch all pages fresh | false | -| `--resume` | - | Resume interrupted crawl | false | -| `--verbose` | `-v` | Verbose output | false | -| `--debug` | - | Debug logging | false | -| `--trace` | - | Enable performance tracing | false | -| `--project-name ` | `-n ` | Override project name | from config | - -### Coverage Modes - -Choose a coverage mode based on your audit needs: - -| Mode | Default Pages | Behavior | Use Case | -|------|---------------|----------|----------| -| `quick` | 25 | Seed + sitemaps only, no link discovery | CI checks, fast health check | -| `surface` | 100 | One sample per URL pattern | General audits (default) | -| `full` | 500 | Crawl everything up to limit | Deep analysis | - -**Surface mode is smart** - it detects URL patterns like `/blog/{slug}` or `/products/{id}` and only crawls one sample per pattern. This makes it efficient for sites with many similar pages (blogs, e-commerce). - -```bash -# Quick health check (25 pages, no link discovery) -squirrel audit https://example.com -C quick --format llm - -# Default surface audit (100 pages, pattern sampling) -squirrel audit https://example.com --format llm - -# Full comprehensive audit (500 pages) -squirrel audit https://example.com -C full --format llm - -# Override page limit for any mode -squirrel audit https://example.com -C surface -m 200 --format llm -``` - -**When to use each mode:** -- `quick`: CI pipelines, daily health checks, monitoring -- `surface`: Most audits - covers unique templates efficiently -- `full`: Before launches, comprehensive analysis, deep dives - -### Report Command Options - -| Option | Alias | Description | -|--------|-------|-------------| -| `--list` | `-l` | List recent audits | -| `--severity ` | - | Filter by severity: error, warning, all | -| `--category ` | - | Filter by categories (comma-separated) | -| `--format ` | `-f ` | Output format: console, text, json, html, markdown, xml, llm | -| `--output ` | `-o ` | Output file path | -| `--input ` | `-i ` | Load from JSON file (fallback mode) | - -### Config Subcommands - -| Command | Description | -|---------|-------------| -| `config show` | Show current config | -| `config set ` | Set config value | -| `config path` | Show config file path | -| `config validate` | Validate config file | - -### Other Commands - -| Command | Description | -|---------|-------------| -| `squirrel feedback` | Send feedback to squirrelscan team | -| `squirrel skills install` | Install Claude Code skill | -| `squirrel skills update` | Update Claude Code skill | - -### Self Commands - -Self-management commands under `squirrel self`: - -| Command | Description | -|---------|-------------| -| `self install` | Bootstrap local installation | -| `self update` | Check and apply updates | -| `self completion` | Generate shell completions | -| `self doctor` | Run health checks | -| `self version` | Show version information | -| `self settings` | Manage CLI settings | -| `self uninstall` | Remove squirrel from the system | - -## Output Formats - -### Console Output (default) - -The `audit` command shows human-readable console output by default with colored output and progress indicators. - -### LLM Format - -To get LLM-optimized output, use the `report` command with `--format llm`: - -```bash -squirrel report --format llm -``` - -The LLM format is a compact XML/text hybrid optimized for token efficiency (40% smaller than verbose XML): - -- **Summary**: Overall health score and key metrics -- **Issues by Category**: Grouped by audit rule category (core SEO, technical, content, security) -- **Broken Links**: List of broken external and internal links -- **Recommendations**: Prioritized action items with fix suggestions - -See [OUTPUT-FORMAT.md](references/OUTPUT-FORMAT.md) for detailed format specification. - -## Examples - -### Example 1: Quick Site Audit with LLM Output - -```bash -# User asks: "Check squirrelscan.com for SEO issues" -squirrel audit https://squirrelscan.com --format llm -``` - -### Example 2: Deep Audit for Large Site - -```bash -# User asks: "Do a thorough audit of my blog with up to 500 pages" -squirrel audit https://myblog.com --max-pages 500 --format llm -``` - -### Example 3: Fresh Audit After Changes - -```bash -# User asks: "Re-audit the site and ignore cached results" -squirrel audit https://example.com --refresh --format llm -``` - -### Example 4: Two-Step Workflow (Reuse Previous Audit) - -```bash -# First run an audit -squirrel audit https://example.com -# Note the audit ID from output (e.g., "a1b2c3d4") - -# Later, export in different format -squirrel report a1b2c3d4 --format llm -``` - -## Output - -On completion give the user a summary of all of the changes you made. - -## Troubleshooting - -### squirrel command not found - -If you see this error, squirrel is not installed or not in your PATH. - -**Solution:** -1. Install squirrel: [squirrelscan.com/download](https://squirrelscan.com/download) -2. Ensure `~/.local/bin` is in PATH -3. Verify: `squirrel --version` - -### Permission denied - -If squirrel is not executable, ensure the binary has execute permissions. Reinstalling from [squirrelscan.com/download](https://squirrelscan.com/download) will fix this. - -### Crawl timeout or slow performance - -For very large sites, the audit may take several minutes. Use `--verbose` to see progress: - -```bash -squirrel audit https://example.com --format llm --verbose -``` - -### Invalid URL - -Ensure the URL includes the protocol (http:// or https://): - -```bash -# ✗ Wrong -squirrel audit example.com - -# ✓ Correct -squirrel audit https://example.com -``` - -## How It Works - -1. **Crawl**: Discovers and fetches pages starting from the base URL -2. **Analyze**: Runs audit rules on each page -3. **External Links**: Checks external links for availability -4. **Report**: Generates LLM-optimized report with findings - -The audit is stored in a local database and can be retrieved later with `squirrel report` commands. - -## Additional Resources - -- **Output Format Reference**: [OUTPUT-FORMAT.md](references/OUTPUT-FORMAT.md) -- **squirrelscan Documentation**: https://docs.squirrelscan.com -- **CLI Help**: `squirrel audit --help` diff --git a/skill_index.md b/skill_index.md new file mode 100644 index 00000000..9658b702 --- /dev/null +++ b/skill_index.md @@ -0,0 +1,562 @@ +| Skill | Description | +|---|---| +| `1password` | name: 1password | +| `ab-test-setup` | name: ab-test-setup | +| `academic-deep-research` | name: academic-deep-research | +| `adclaw` | name: adclaw | +| `add-educational-comments` | name: add-educational-comments | +| `adwords` | name: adwords | +| `agent-autonomy-kit` | name: agent-autonomy-kit | +| `agent-browser` | name: Agent Browser | +| `agent-browser-clawdbot` | name: agent-browser | +| `agent-designer` | name: "agent-designer" | +| `agent-governance` | name: agent-governance | +| `agent-memory` | Persistent memory system for AI agents. Remember facts, learn from experience, and track entities ac | +| `agent-reach` | name: agent-reach | +| `agent-team-orchestration` | name: agent-team-orchestration | +| `agent-workflow-designer` | name: "agent-workflow-designer" | +| `agentcreate` | name: agentCreate | +| `agentic-eval` | name: agentic-eval | +| `agentmail` | name: agentmail | +| `agile-product-owner` | name: "agile-product-owner" | +| `ai-humanizer` | name: humanizer | +| `ai-model-router` | name: ai-model-router | +| `ai-model-router-v2` | name: ai-model-router | +| `ai-news-aggregator-sl` | name: ai-news-aggregator-sl | +| `ai-ppt-generator` | name: ai-ppt-generator | +| `ai-prompt-engineering-safety-review` | name: ai-prompt-engineering-safety-review | +| `ai-prompt-generator` | 专业 AI 提示词生成工具,帮助用户创建高效、精准的 AI 提示词。内置多种框架和模板,让 AI 输出质量提升 10 倍。 | +| `ai-task-hub` | name: ai-task-hub | +| `ai-travel` | name: ai-travel | +| `ai-web-automation` | 自动化 Web 任务执行服务。 | +| `alchemy-openapi-skill` | name: alchemy-openapi-skill | +| `algorithmic-art` | name: algorithmic-art | +| `amazon-price-tracker` | 实时监控亚马逊商品价格,设置降价提醒,追踪历史价格曲线。帮助买家低价购入,卖家竞品监控。 | +| `analytics-tracking` | name: analytics-tracking | +| `answeroverflow` | name: answeroverflow | +| `api-design-principles` | name: api-design-principles | +| `api-design-reviewer` | name: "api-design-reviewer" | +| `api-gateway` | name: api-gateway | +| `api-test-suite-builder` | name: "api-test-suite-builder" | +| `app-store-optimization` | name: "app-store-optimization" | +| `apple-appstore-reviewer` | name: apple-appstore-reviewer | +| `apple-notes` | name: apple-notes | +| `apple-reminders` | name: apple-reminders | +| `architecture-blueprint-generator` | name: architecture-blueprint-generator | +| `architecture-patterns` | name: architecture-patterns | +| `atlassian-admin` | name: "atlassian-admin" | +| `atlassian-templates` | name: "atlassian-templates" | +| `audit-website` | name: audit-website | +| `auto-memory-pro` | name: auto-memory-pro | +| `auto-updater` | name: auto-updater | +| `automation-workflows` | name: automation-workflows | +| `autoresearch-agent` | name: "autoresearch-agent" | +| `aws-solution-architect` | name: "aws-solution-architect" | +| `baidu-search` | name: baidu-search | +| `bankofbots` | name: bankofbots | +| `basedagents` | name: basedagents | +| `bear-notes` | name: bear-notes | +| `better-auth-best-practices` | name: better-auth-best-practices | +| `bird` | name: bird | +| `blogwatcher` | name: blogwatcher | +| `blucli` | name: blucli | +| `bluebubbles` | name: bluebubbles | +| `board-deck-builder` | name: "board-deck-builder" | +| `board-meeting` | name: "board-meeting" | +| `boost-prompt` | name: boost-prompt | +| `brainstorming` | name: brainstorming | +| `brand-guidelines` | name: brand-guidelines | +| `brave-search` | name: brave-search | +| `breakdown-feature-implementation` | name: breakdown-feature-implementation | +| `browser` | This skill uses a headless browser (Puppeteer) to render web pages and extract clean, readable conte | +| `browser-use` | name: browser-use | +| `business-growth` | name: "business-growth-skills" | +| `byterover` | name: byterover | +| `c-level-advisor` | name: "c-level-advisor" | +| `c-suite-agent-protocol` | name: agent-protocol | +| `c-suite-competitive-intel` | name: competitive-intel | +| `c-suite-founder-coach` | name: founder-coach | +| `caldav-calendar` | name: caldav-calendar | +| `calendar` | name: calendar | +| `campaign-analytics` | name: "campaign-analytics" | +| `camsnap` | name: camsnap | +| `canvas` | Display HTML content on connected OpenClaw nodes (Mac app, iOS, Android). | +| `canvas-design` | name: canvas-design | +| `capa-officer` | name: "capa-officer" | +| `capability-evolver` | name: capability-evolver | +| `ceo-advisor` | name: "ceo-advisor" | +| `cfo-advisor` | name: "cfo-advisor" | +| `chainbase-openapi-skill` | name: chainbase-openapi-skill | +| `change-management` | name: "change-management" | +| `changelog-curator` | name: changelog-curator | +| `changelog-generator` | name: "changelog-generator" | +| `chief-of-staff` | name: "chief-of-staff" | +| `chro-advisor` | name: "chro-advisor" | +| `chrome-devtools` | name: chrome-devtools | +| `ci-cd-pipeline-builder` | name: "ci-cd-pipeline-builder" | +| `ciso-advisor` | name: "ciso-advisor" | +| `citedy-content-ingestion` | name: citedy-content-ingestion | +| `citedy-content-writer` | name: citedy-content-writer | +| `citedy-lead-magnets` | name: citedy-lead-magnets | +| `citedy-trend-scout` | name: citedy-trend-scout | +| `citedy-video-shorts` | name: citedy-video-shorts | +| `citrea-claw-skill` | name: citrea-claw-skill | +| `clankers-world` | name: "Clanker's World" | +| `clawdbot-filesystem` | name: filesystem | +| `clawddocs` | name: clawddocs | +| `clawdhub` | name: clawdhub | +| `clawsec` | | +| `clean-content-fetch` | name: clean-content-fetch | +| `clipboard-knowledge-capture` | name: clipboard-knowledge-capture | +| `cmo-advisor` | name: "cmo-advisor" | +| `code` | name: Code | +| `code-exemplars-blueprint-generator` | name: code-exemplars-blueprint-generator | +| `code-review` | name: code-review | +| `code-reviewer-2` | name: code-reviewer | +| `code-to-prd` | Name: code-to-prd | +| `codebase-onboarding` | name: "codebase-onboarding" | +| `coding` | name: Coding | +| `coding-agent` | name: coding-agent | +| `coingecko-openapi-skill` | name: coingecko-openapi-skill | +| `communication-playbook` | name: communication-playbook | +| `compaction-ui-enhancements` | name: compaction-ui | +| `company-os` | name: "company-os" | +| `competitive-teardown` | name: "competitive-teardown" | +| `competitor-alternatives` | name: competitor-alternatives | +| `computer-use` | name: computer-use | +| `confluence-expert` | name: "confluence-expert" | +| `content-creator` | name: "content-creator" | +| `content-humanizer` | name: "content-humanizer" | +| `content-production` | name: "content-production" | +| `content-strategy` | name: content-strategy | +| `context-engine` | name: "context-engine" | +| `contract-and-proposal-writer` | name: "contract-and-proposal-writer" | +| `coo-advisor` | name: "coo-advisor" | +| `copy-editing` | name: copy-editing | +| `copywriting` | name: copywriting | +| `cpo-advisor` | name: "cpo-advisor" | +| `create-auth-skill` | name: create-auth-skill | +| `create-cli` | name: create-cli | +| `creative-toolkit` | name: "AI Image Generation & Editor — Nanobanana, GPT Image, ComfyUI" | +| `cro-advisor` | name: "cro-advisor" | +| `cron-mastery` | name: cron-mastery | +| `cs-ab-test-setup` | name: "ab-test-setup" | +| `cs-ad-creative` | name: "ad-creative" | +| `cs-agent-protocol` | name: "agent-protocol" | +| `cs-ai-seo` | name: "ai-seo" | +| `cs-analytics-tracking` | name: "analytics-tracking" | +| `cs-brand-guidelines` | name: "brand-guidelines" | +| `cs-churn-prevention` | name: "churn-prevention" | +| `cs-code-reviewer` | name: "code-reviewer" | +| `cs-cold-email` | name: "cold-email" | +| `cs-competitive-intel` | name: "competitive-intel" | +| `cs-competitor-alternatives` | name: "competitor-alternatives" | +| `cs-content-strategy` | name: "content-strategy" | +| `cs-copywriting` | name: "copywriting" | +| `cs-financial-analyst` | name: "financial-analyst" | +| `cs-founder-coach` | name: "founder-coach" | +| `cs-google-workspace-cli` | name: "google-workspace-cli" | +| `cs-landing-page-generator` | name: "landing-page-generator" | +| `cs-onboard` | name: "cs-onboard" | +| `cs-performance-profiler` | name: "performance-profiler" | +| `cs-playwright-pro` | name: "playwright-pro" | +| `cs-pricing-strategy` | name: "pricing-strategy" | +| `cs-pw` | name: "playwright-pro" | +| `cs-schema-markup` | name: "schema-markup" | +| `cs-self-improving-agent` | name: "self-improving-agent" | +| `cs-seo-audit` | name: "seo-audit" | +| `cs-skill-security-auditor` | name: "skill-security-auditor" | +| `cs-social-content` | name: "social-content" | +| `cs-social-media-manager` | name: "social-media-manager" | +| `cto-advisor` | name: "cto-advisor" | +| `culture-architect` | name: "culture-architect" | +| `customer-success-manager` | name: "customer-success-manager" | +| `data-analysis` | name: Data Analysis | +| `data-analyst` | name: data-analyst | +| `database-admin` | name: database-admin | +| `database-designer` | name: "database-designer" | +| `database-schema-designer` | name: "database-schema-designer" | +| `ddg-web-search` | name: ddg-search | +| `debug-pro` | Systematic debugging methodology and language-specific debugging commands. | +| `decision-logger` | name: "decision-logger" | +| `deep-research-pro` | name: deep-research-pro | +| `dependency-auditor` | name: "dependency-auditor" | +| `desearch-web-search` | name: desearch-web-search | +| `desktop-control` | description: Advanced desktop automation with mouse, keyboard, and screen control | +| `discord` | name: discord | +| `dispatching-parallel-agents` | name: dispatching-parallel-agents | +| `doc-coauthoring` | name: doc-coauthoring | +| `docker-development` | name: "docker-development" | +| `docker-essentials` | name: docker-essentials | +| `document-parser` | 高精度文档解析技能,从 PDF、图片、Word 文档中提取结构化数据。 | +| `docx` | name: docx | +| `domain-dns-ops` | name: domain-dns-ops | +| `duckduckgo-search` | name: duckduckgo-search | +| `eastmoney-financial-data-1-0-2` | name: eastmoney_financial_data | +| `eastmoney-financial-search-1-0-2` | name: eastmoney_financial_search | +| `ebay-product-research` | 专业 eBay 选品分析工具,帮助卖家发现高利润、低竞争的产品。分析销量、价格趋势、竞争程度、利润空间,提供数据驱动的选品建议。 | +| `edge-tts` | name: edge-tts | +| `eightctl` | name: eightctl | +| `elite-longterm-memory` | name: elite-longterm-memory | +| `email-sequence` | name: email-sequence | +| `email-template-builder` | name: "email-template-builder" | +| `env-secrets-manager` | name: "env-secrets-manager" | +| `epic-design` | name: epic-design | +| `erpclaw` | name: erpclaw | +| `evomap` | name: evomap | +| `exa-web-search-free` | name: exa-web-search-free | +| `excel-xlsx` | name: Excel / XLSX | +| `executing-plans` | name: executing-plans | +| `executive-mentor` | name: "executive-mentor" | +| `experiment-designer` | name: experiment-designer | +| `expo-api-routes` | name: expo-api-routes | +| `expo-building-native-ui` | name: building-native-ui | +| `expo-cicd-workflows` | name: expo-cicd-workflows | +| `expo-deployment` | name: expo-deployment | +| `expo-dev-client` | name: expo-dev-client | +| `expo-native-data-fetching` | name: native-data-fetching | +| `expo-tailwind-setup` | name: expo-tailwind-setup | +| `expo-ui-jetpack-compose` | name: Expo UI Jetpack Compose | +| `expo-ui-swift-ui` | name: Expo UI SwiftUI | +| `expo-use-dom` | name: use-dom | +| `fda-consultant-specialist` | name: "fda-consultant-specialist" | +| `feishu-doc` | name: feishu-doc | +| `feishu-doc-collab` | name: feishu-doc-collab | +| `feishu-evolver-wrapper` | name: feishu-evolver-wrapper | +| `file-search` | name: file-search | +| `filesystem` | name: filesystem | +| `find-skills` | name: find-skills | +| `finishing-a-development-branch` | name: finishing-a-development-branch | +| `firecrawl` | name: firecrawl | +| `firecrawl-search` | name: firecrawl | +| `food-order` | name: food-order | +| `form-cro` | name: form-cro | +| `free-ride` | name: freeride | +| `free-tool-strategy` | name: free-tool-strategy | +| `frontend-design` | name: frontend-design | +| `frontend-design-ultimate` | name: frontend-design-ultimate | +| `gcalcli-calendar` | name: gcalcli-calendar | +| `gdpr-dsgvo-expert` | name: "gdpr-dsgvo-expert" | +| `gemini` | name: gemini | +| `gemini-browser` | name: Gemini Browser | +| `gifgrep` | name: gifgrep | +| `git` | name: Git (Essentials + Workflows + Advanced) | +| `git-commit` | name: git-commit | +| `git-essentials` | name: git-essentials | +| `git-worktree-manager` | name: "git-worktree-manager" | +| `github` | name: github | +| `gmail` | name: gmail | +| `go-install` | Content-Disposition: form-data; name="file"; filename="SKILL.md" | +| `go-install-zh` | Content-Disposition: form-data; name="file"; filename="SKILL.md" | +| `gog` | name: gog | +| `gogcli` | name: gogcli | +| `google-calendar` | name: google-calendar | +| `google-search` | name: google-search | +| `goplaces` | name: goplaces | +| `haodf` | name: haodf | +| `health-score-pro` | name: health-management | +| `healthcheck` | name: healthcheck | +| `helm-chart-builder` | name: "helm-chart-builder" | +| `hengheng-system-time` | name: system-time | +| `himalaya` | name: himalaya | +| `home-assistant` | name: home-assistant | +| `html-slide-creator` | name: slide-creator | +| `humanize-ai-text` | name: humanize-ai-text | +| `humanizer` | name: humanizer | +| `image-generate` | name: image-generate | +| `imap-smtp-email` | name: imap-smtp-email | +| `imsg` | name: imsg | +| `incident-commander` | name: "incident-commander" | +| `information-security-manager-iso27001` | name: "information-security-manager-iso27001" | +| `instruments-profiling` | name: instruments-profiling | +| `internal-comms` | name: internal-comms | +| `internal-narrative` | name: "internal-narrative" | +| `interview-system-designer` | name: "interview-system-designer" | +| `intl-expansion` | name: "intl-expansion" | +| `isms-audit-expert` | name: "isms-audit-expert" | +| `jike-publisher` | name: jike-publisher | +| `jira-expert` | name: "jira-expert" | +| `last30days` | name: last30days | +| `launch-strategy` | name: launch-strategy | +| `linear` | name: linear | +| `linkedin` | name: linkedin | +| `local-places` | name: local-places | +| `ltx-video` | name: ltx-video | +| `ma-playbook` | name: "ma-playbook" | +| `main-image-editor` | name: main-image-editor | +| `markdown-converter` | name: markdown-converter | +| `marketing-context` | name: "marketing-context" | +| `marketing-demand-acquisition` | name: "marketing-demand-acquisition" | +| `marketing-ideas` | name: marketing-ideas | +| `marketing-mode` | name: marketing-mode | +| `marketing-ops` | name: "marketing-ops" | +| `marketing-psychology` | name: marketing-psychology | +| `marketing-skills` | name: marketing-skills | +| `marketing-strategy-pmm` | name: "marketing-strategy-pmm" | +| `master-skills` | name: bagman | +| `mcp-builder` | name: mcp-builder | +| `mcp-server-builder` | name: "mcp-server-builder" | +| `mcporter` | name: mcporter | +| `mdr-745-specialist` | name: "mdr-745-specialist" | +| `media-generation` | name: media-generation | +| `memory-hygiene` | name: memory-hygiene | +| `memory-manager` | name: memory-manager | +| `memory-setup` | name: memory-setup | +| `microservices-patterns` | name: microservices-patterns | +| `microsoft-excel` | name: microsoft-excel | +| `microsoft-skill-creator` | name: microsoft-skill-creator | +| `migration-architect` | name: "migration-architect" | +| `mindkeeper` | name: mindkeeper | +| `miniade-agent-lifecycle-manager` | name: agent-lifecycle-manager | +| `model-usage` | name: model-usage | +| `modern-javascript-patterns` | name: modern-javascript-patterns | +| `moltbook-interact` | name: moltbook | +| `monorepo-navigator` | name: "monorepo-navigator" | +| `moralis-openapi-skill` | name: moralis-openapi-skill | +| `ms365-tenant-manager` | name: "ms365-tenant-manager" | +| `multi-search-engine` | name: "multi-search-engine" | +| `n8n` | name: n8n | +| `n8n-workflow-automation` | name: n8n-workflow-automation | +| `nano-banana-pro` | name: nano-banana-pro | +| `nano-pdf` | name: nano-pdf | +| `native-app-performance` | name: native-app-performance | +| `new-visitor-cold-start` | name: new-visitor-cold-start | +| `news-summary` | name: news-summary | +| `next-best-practices` | name: next-best-practices | +| `next-cache-components` | name: next-cache-components | +| `nextjs-app-router-patterns` | name: nextjs-app-router-patterns | +| `nidhov01-agent-browser` | name: Agent Browser | +| `nidhov01-find-skills` | name: find-skills | +| `nidhov01-github` | name: github | +| `nidhov01-notion` | name: notion | +| `nidhov01-proactive-agent` | name: proactive-agent | +| `nodejs-backend-patterns` | name: nodejs-backend-patterns | +| `notion` | name: notion | +| `nuxt` | name: nuxt | +| `observability-designer` | name: "observability-designer" | +| `obsidian` | name: obsidian | +| `offer-positioning-auditor` | name: offer-positioning-auditor | +| `onboarding-cro` | name: onboarding-cro | +| `openai-codex-multi-oauth` | name: openai-codex-multi-oauth | +| `openai-image-gen` | name: openai-image-gen | +| `openai-whisper` | name: openai-whisper | +| `openai-whisper-api` | name: openai-whisper-api | +| `openclaw-backup` | name: openclaw-backup | +| `openclaw-guardian` | name: openclaw-guardian | +| `openclaw-skill-vetter` | name: skill-vetter | +| `openclaw-tavily-search` | name: tavily-search | +| `openclaw-todoist` | name: todoist | +| `opencode-controller` | name: opencode-controller | +| `openhue` | name: openhue | +| `oracle` | name: oracle | +| `ordercli` | name: ordercli | +| `org-health-diagnostic` | name: "org-health-diagnostic" | +| `outlook` | name: outlook | +| `page-cro` | name: page-cro | +| `paid-ads` | name: paid-ads | +| `para-second-brain` | name: para-second-brain | +| `partnerships-ecosystem` | name: partnerships-ecosystem | +| `paywall-upgrade-cro` | name: "paywall-upgrade-cro" | +| `pdf` | name: pdf | +| `pdf-extract` | name: pdf-extract | +| `pdf-text-extractor` | name: pdf-text-extractor | +| `peekaboo` | name: peekaboo | +| `perplexity` | name: perplexity | +| `personal-finish-notifier` | name: personal-finish-notifier | +| `pinia` | name: pinia | +| `playwright` | name: Playwright (Automation + MCP + Scraper) | +| `playwright-mcp` | name: playwright-mcp | +| `playwright-pro` | name: "playwright-pro" | +| `pnpm` | name: pnpm | +| `popup-cro` | name: popup-cro | +| `postgresql-table-design` | name: postgresql-table-design | +| `ppt-generator` | name: ppt-generator | +| `pptx` | name: pptx | +| `pr-review-expert` | name: "pr-review-expert" | +| `pricing-strategy` | name: pricing-strategy | +| `proactive-agent` | name: proactive-agent | +| `proactive-agent-lite` | name: proactive-agent-lite | +| `product-analytics` | name: product-analytics | +| `product-dev-ops-package` | name: product-dev-ops-team | +| `product-discovery` | name: product-discovery | +| `product-manager-toolkit` | name: "product-manager-toolkit" | +| `product-marketing-context` | name: product-marketing-context | +| `product-strategist` | name: "product-strategist" | +| `productivity` | name: Productivity | +| `programmatic-seo` | name: programmatic-seo | +| `prompt-engineer-toolkit` | name: "prompt-engineer-toolkit" | +| `prompt-engineering-expert` | name: prompt-engineering-expert | +| `prompt-engineering-patterns` | name: prompt-engineering-patterns | +| `python-design-patterns` | name: python-design-patterns | +| `python-performance-optimization` | name: python-performance-optimization | +| `python-testing-patterns` | name: python-testing-patterns | +| `qmd` | name: qmd | +| `qms-audit-expert` | name: "qms-audit-expert" | +| `quality-documentation-manager` | name: "quality-documentation-manager" | +| `quality-manager-qmr` | name: "quality-manager-qmr" | +| `quality-manager-qms-iso13485` | name: "quality-manager-qms-iso13485" | +| `rag-architect` | name: "rag-architect" | +| `rag-implementation` | name: rag-implementation | +| `react-doctor` | name: react-doctor | +| `react-native-best-practices` | name: react-native-best-practices | +| `react-state-management` | name: react-state-management | +| `readgzh` | name: readgzh | +| `receiving-code-review` | name: receiving-code-review | +| `reddit` | name: reddit | +| `reddit-readonly` | name: reddit-readonly | +| `referral-program` | name: referral-program | +| `regulatory-affairs-head` | name: "regulatory-affairs-head" | +| `release-manager` | name: "release-manager" | +| `remembering-conversations` | name: remembering-conversations | +| `requesting-code-review` | name: requesting-code-review | +| `research-summarizer` | name: "research-summarizer" | +| `responsive-design` | name: responsive-design | +| `revenue-operations` | name: "revenue-operations" | +| `risk-management-specialist` | name: "risk-management-specialist" | +| `roadmap-communicator` | name: roadmap-communicator | +| `runbook-generator` | name: "runbook-generator" | +| `runtime-sentinel` | name: runtime-sentinel | +| `rustchain-mcp` | MCP server giving AI agents access to the RustChain Proof-of-Antiquity blockchain, BoTTube AI-native | +| `saas-metrics-coach` | name: saas-metrics-coach | +| `saas-scaffolder` | name: "saas-scaffolder" | +| `safe-exec` | name: safe-exec | +| `sag` | name: sag | +| `sales-engineer` | name: "sales-engineer" | +| `salesmate` | name: salesmate | +| `scenario-war-room` | name: "scenario-war-room" | +| `scrapling-official` | name: scrapling-official | +| `scrum-master` | name: "scrum-master" | +| `searxng` | name: searxng | +| `security-auditor` | name: security-auditor | +| `seek-and-analyze-video` | name: seek-and-analyze-video | +| `self-improving` | name: Self-Improving Agent (Proactive Self-Reflection) | +| `self-reflection` | name: self-reflection | +| `senior-architect` | name: "senior-architect" | +| `senior-backend` | name: "senior-backend" | +| `senior-computer-vision` | name: "senior-computer-vision" | +| `senior-data-engineer` | name: "senior-data-engineer" | +| `senior-data-scientist` | name: "senior-data-scientist" | +| `senior-devops` | name: "senior-devops" | +| `senior-frontend` | name: "senior-frontend" | +| `senior-fullstack` | name: "senior-fullstack" | +| `senior-ml-engineer` | name: "senior-ml-engineer" | +| `senior-pm` | name: "senior-pm" | +| `senior-prompt-engineer` | name: "senior-prompt-engineer" | +| `senior-qa` | name: "senior-qa" | +| `senior-secops` | name: "senior-secops" | +| `senior-security` | name: "senior-security" | +| `sentinel-oleg` | name: claw-sentinel | +| `seo-audit` | name: seo-audit | +| `session-logs` | name: session-logs | +| `sglang-diffusion-video` | name: sglang-diffusion-video | +| `shopify-seo-bot` | 自动优化 Shopify 店铺 SEO,包括产品标题、描述、meta 标签、图片 ALT 等。提升 Google 搜索排名,增加自然流量。 | +| `shopify-seo-optimizer` | 专为 Shopify 店铺设计的 SEO 优化工具。优化产品标题、描述、元标签、图片 Alt、URL 结构,提升店铺在 Google 的搜索排名和自然流量。 | +| `signup-flow-cro` | name: signup-flow-cro | +| `site-architecture` | name: "site-architecture" | +| `skill-creator` | name: skill-creator | +| `skill-finder-cn` | name: skill-finder-cn | +| `skill-listing-polisher` | name: skill-listing-polisher | +| `skill-scanner` | name: skill-scanner | +| `skill-tester` | name: "skill-tester" | +| `skill-vetter` | name: skill-vetter | +| `skill-vetting` | name: skill-vetting | +| `slack` | name: slack | +| `slack-gif-creator` | name: slack-gif-creator | +| `slidev` | name: slidev | +| `social-content` | name: social-content | +| `social-media-analyzer` | name: "social-media-analyzer" | +| `songsee` | name: songsee | +| `sonoscli` | name: sonoscli | +| `spotify-player` | name: spotify-player | +| `sql-toolkit` | name: sql-toolkit | +| `stock-analysis` | name: stock-analysis | +| `stock-market-pro` | name: stock-market-pro | +| `stock-watcher` | name: stock-watcher | +| `strategic-alignment` | name: "strategic-alignment" | +| `stripe-integration-expert` | name: "stripe-integration-expert" | +| `subagent-driven-development` | name: subagent-driven-development | +| `summarize` | name: summarize | +| `supabase-postgres-best-practices` | name: supabase-postgres-best-practices | +| `superdesign` | name: frontend-design | +| `swarmclaw` | name: swarmclaw | +| `swift-concurrency-expert` | name: swift-concurrency-expert | +| `swiftui-liquid-glass` | name: swiftui-liquid-glass | +| `swiftui-performance-audit` | name: swiftui-performance-audit | +| `swiftui-view-refactor` | name: swiftui-view-refactor | +| `systematic-debugging` | name: systematic-debugging | +| `tavily` | name: tavily | +| `tavily-search-1-0-0` | name: tavily | +| `tdd-guide` | name: "tdd-guide" | +| `tech-data-playbook` | name: tech-data-playbook | +| `tech-debt-tracker` | name: tech-debt-tracker | +| `tech-stack-evaluator` | name: "tech-stack-evaluator" | +| `telegram` | name: telegram | +| `template-skill` | name: template-skill | +| `terraform-patterns` | name: "terraform-patterns" | +| `test-driven-development` | name: test-driven-development | +| `theme-factory` | name: theme-factory | +| `things-mac` | name: things-mac | +| `ths-advanced-analysis` | name: ths-advanced-analysis | +| `tiktok-viral-predictor` | AI 预测 TikTok 视频爆款潜力,分析热门元素、BGM、标签。提供优化建议,提高视频上推荐概率。 | +| `tmux` | name: tmux | +| `todo-tracker-safe` | name: todo-tracker-safe | +| `todoist` | name: todoist | +| `trader-daily` | name: quant-trader-daily | +| `trello` | name: trello | +| `turborepo` | name: turborepo | +| `turing-pyramid` | name: turing-pyramid | +| `tushare-finance` | name: tushare-finance | +| `typescript-advanced-types` | name: typescript-advanced-types | +| `ui-design-system` | name: "ui-design-system" | +| `ui-ux-pro-max` | name: ui-ux-pro-max | +| `unocss` | name: unocss | +| `upbit-openapi-skill` | name: upbit-openapi-skill | +| `upgrading-expo` | name: upgrading-expo | +| `upgrading-react-native` | name: upgrading-react-native | +| `us-stock-analysis` | name: us-stock-analysis | +| `using-git-worktrees` | name: using-git-worktrees | +| `using-superpowers` | name: using-superpowers | +| `ux-researcher-designer` | name: "ux-researcher-designer" | +| `veadk-skills` | name: veadk-skills | +| `vercel-ai-sdk` | name: ai-sdk | +| `vercel-composition-patterns` | name: vercel-composition-patterns | +| `vercel-react-best-practices` | name: vercel-react-best-practices | +| `verification-before-completion` | name: verification-before-completion | +| `video-frames` | name: video-frames | +| `video-transcript-downloader` | name: video-transcript-downloader | +| `vite` | name: vite | +| `vitepress` | name: vitepress | +| `vitest` | name: vitest | +| `vue` | name: vue | +| `vue-best-practices` | name: vue-best-practices | +| `vue-best-practices-hyf0` | name: vue-best-practices | +| `vue-debug-guides` | name: vue-debug-guides | +| `vue-jsx-best-practices` | name: vue-jsx-best-practices | +| `vue-pinia-best-practices` | name: vue-pinia-best-practices | +| `vue-router-best-practices` | name: vue-router-best-practices | +| `vue-router-best-practices-hyf0` | name: vue-router-best-practices | +| `vue-testing-best-practices` | name: vue-testing-best-practices | +| `vue-testing-best-practices-hyf0` | name: vue-testing-best-practices | +| `wacli` | name: wacli | +| `weather` | name: weather | +| `web-artifacts-builder` | name: web-artifacts-builder | +| `web-component-design` | name: web-component-design | +| `web-design-guidelines` | name: web-design-guidelines | +| `web-search-plus` | name: web-search-plus | +| `webapp-testing` | name: webapp-testing | +| `weibo-trending-bot` | 实时监控微博热搜榜,追踪热点话题、明星八卦、社会新闻。自动生成蹭热点文案。 | +| `widget` | name: widget | +| `word-docx` | name: Word / Docx | +| `writing-plans` | name: writing-plans | +| `writing-skills` | name: writing-skills | +| `wyckoff-a-share` | name: akshare-online-alpha | +| `x-twitter` | name: twitter-openclaw | +| `x-twitter-growth` | name: "x-twitter-growth" | +| `xiaohongshu-mcp` | name: xiaohongshu-mcp | +| `xlsx` | name: xlsx | +| `xurl` | name: xurl | +| `yahoo-finance` | name: yahoo-finance | +| `youtube-api-skill` | name: youtube | +| `youtube-auto-captions` | 自动为 YouTube 视频生成字幕,支持多语言翻译、时间轴校准。提升视频可访问性和 SEO。 | +| `youtube-transcript` | name: youtube-transcript | +| `youtube-watcher` | name: youtube-watcher | diff --git a/skills/agent-designer/README.md b/skills/agent-designer/README.md new file mode 100644 index 00000000..5a023e71 --- /dev/null +++ b/skills/agent-designer/README.md @@ -0,0 +1,430 @@ +# Agent Designer - Multi-Agent System Architecture Toolkit + +**Tier:** POWERFUL +**Category:** Engineering +**Tags:** AI agents, architecture, system design, orchestration, multi-agent systems + +A comprehensive toolkit for designing, architecting, and evaluating multi-agent systems. Provides structured approaches to agent architecture patterns, tool design principles, communication strategies, and performance evaluation frameworks. + +## Overview + +The Agent Designer skill includes three core components: + +1. **Agent Planner** (`agent_planner.py`) - Designs multi-agent system architectures +2. **Tool Schema Generator** (`tool_schema_generator.py`) - Creates structured tool schemas +3. **Agent Evaluator** (`agent_evaluator.py`) - Evaluates system performance and identifies optimizations + +## Quick Start + +### 1. Design a Multi-Agent Architecture + +```bash +# Use sample requirements or create your own +python agent_planner.py assets/sample_system_requirements.json -o my_architecture + +# This generates: +# - my_architecture.json (complete architecture) +# - my_architecture_diagram.mmd (Mermaid diagram) +# - my_architecture_roadmap.json (implementation plan) +``` + +### 2. Generate Tool Schemas + +```bash +# Use sample tool descriptions or create your own +python tool_schema_generator.py assets/sample_tool_descriptions.json -o my_tools + +# This generates: +# - my_tools.json (complete schemas) +# - my_tools_openai.json (OpenAI format) +# - my_tools_anthropic.json (Anthropic format) +# - my_tools_validation.json (validation rules) +# - my_tools_examples.json (usage examples) +``` + +### 3. Evaluate System Performance + +```bash +# Use sample execution logs or your own +python agent_evaluator.py assets/sample_execution_logs.json -o evaluation + +# This generates: +# - evaluation.json (complete report) +# - evaluation_summary.json (executive summary) +# - evaluation_recommendations.json (optimization suggestions) +# - evaluation_errors.json (error analysis) +``` + +## Detailed Usage + +### Agent Planner + +The Agent Planner designs multi-agent architectures based on system requirements. + +#### Input Format + +Create a JSON file with system requirements: + +```json +{ + "goal": "Your system's primary objective", + "description": "Detailed system description", + "tasks": ["List", "of", "required", "tasks"], + "constraints": { + "max_response_time": 30000, + "budget_per_task": 1.0, + "quality_threshold": 0.9 + }, + "team_size": 6, + "performance_requirements": { + "high_throughput": true, + "fault_tolerance": true, + "low_latency": false + }, + "safety_requirements": [ + "Input validation and sanitization", + "Output content filtering" + ] +} +``` + +#### Command Line Options + +```bash +python agent_planner.py [OPTIONS] + +Options: + -o, --output PREFIX Output file prefix (default: agent_architecture) + --format FORMAT Output format: json, both (default: both) +``` + +#### Output Files + +- **Architecture JSON**: Complete system design with agents, communication topology, and scaling strategy +- **Mermaid Diagram**: Visual representation of the agent architecture +- **Implementation Roadmap**: Phased implementation plan with timelines and risks + +#### Architecture Patterns + +The planner automatically selects from these patterns based on requirements: + +- **Single Agent**: Simple, focused tasks (1 agent) +- **Supervisor**: Hierarchical delegation (2-8 agents) +- **Swarm**: Peer-to-peer collaboration (3-20 agents) +- **Hierarchical**: Multi-level management (5-50 agents) +- **Pipeline**: Sequential processing (3-15 agents) + +### Tool Schema Generator + +Generates structured tool schemas compatible with OpenAI and Anthropic formats. + +#### Input Format + +Create a JSON file with tool descriptions: + +```json +{ + "tools": [ + { + "name": "tool_name", + "purpose": "What the tool does", + "category": "Tool category (search, data, api, etc.)", + "inputs": [ + { + "name": "parameter_name", + "type": "string", + "description": "Parameter description", + "required": true, + "examples": ["example1", "example2"] + } + ], + "outputs": [ + { + "name": "result_field", + "type": "object", + "description": "Output description" + } + ], + "error_conditions": ["List of possible errors"], + "side_effects": ["List of side effects"], + "idempotent": true, + "rate_limits": { + "requests_per_minute": 60 + } + } + ] +} +``` + +#### Command Line Options + +```bash +python tool_schema_generator.py [OPTIONS] + +Options: + -o, --output PREFIX Output file prefix (default: tool_schemas) + --format FORMAT Output format: json, both (default: both) + --validate Validate generated schemas +``` + +#### Output Files + +- **Complete Schemas**: All schemas with validation and examples +- **OpenAI Format**: Schemas compatible with OpenAI function calling +- **Anthropic Format**: Schemas compatible with Anthropic tool use +- **Validation Rules**: Input validation specifications +- **Usage Examples**: Example calls and responses + +#### Schema Features + +- **Input Validation**: Comprehensive parameter validation rules +- **Error Handling**: Structured error response formats +- **Rate Limiting**: Configurable rate limit specifications +- **Documentation**: Auto-generated usage examples +- **Security**: Built-in security considerations + +### Agent Evaluator + +Analyzes agent execution logs to identify performance issues and optimization opportunities. + +#### Input Format + +Create a JSON file with execution logs: + +```json +{ + "execution_logs": [ + { + "task_id": "unique_task_identifier", + "agent_id": "agent_identifier", + "task_type": "task_category", + "start_time": "2024-01-15T09:00:00Z", + "end_time": "2024-01-15T09:02:34Z", + "duration_ms": 154000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2300, + "success": true + } + ], + "results": { + "summary": "Task results", + "quality_score": 0.92 + }, + "tokens_used": { + "input_tokens": 1250, + "output_tokens": 2800, + "total_tokens": 4050 + }, + "cost_usd": 0.081, + "error_details": null, + "tools_used": ["web_search"], + "retry_count": 0 + } + ] +} +``` + +#### Command Line Options + +```bash +python agent_evaluator.py [OPTIONS] + +Options: + -o, --output PREFIX Output file prefix (default: evaluation_report) + --format FORMAT Output format: json, both (default: both) + --detailed Include detailed analysis in output +``` + +#### Output Files + +- **Complete Report**: Comprehensive performance analysis +- **Executive Summary**: High-level metrics and health assessment +- **Optimization Recommendations**: Prioritized improvement suggestions +- **Error Analysis**: Detailed error patterns and solutions + +#### Evaluation Metrics + +**Performance Metrics**: +- Task success rate and completion times +- Token usage and cost efficiency +- Error rates and retry patterns +- Throughput and latency distributions + +**System Health**: +- Overall health score (poor/fair/good/excellent) +- SLA compliance tracking +- Resource utilization analysis +- Trend identification + +**Bottleneck Analysis**: +- Agent performance bottlenecks +- Tool usage inefficiencies +- Communication overhead +- Resource constraints + +## Architecture Patterns Guide + +### When to Use Each Pattern + +#### Single Agent +- **Best for**: Simple, focused tasks with clear boundaries +- **Team size**: 1 agent +- **Complexity**: Low +- **Examples**: Personal assistant, document summarizer, simple automation + +#### Supervisor +- **Best for**: Hierarchical task decomposition with quality control +- **Team size**: 2-8 agents +- **Complexity**: Medium +- **Examples**: Research coordinator with specialists, content review workflow + +#### Swarm +- **Best for**: Distributed problem solving with high fault tolerance +- **Team size**: 3-20 agents +- **Complexity**: High +- **Examples**: Parallel data processing, distributed research, competitive analysis + +#### Hierarchical +- **Best for**: Large-scale operations with organizational structure +- **Team size**: 5-50 agents +- **Complexity**: Very High +- **Examples**: Enterprise workflows, complex business processes + +#### Pipeline +- **Best for**: Sequential processing with specialized stages +- **Team size**: 3-15 agents +- **Complexity**: Medium +- **Examples**: Data ETL pipelines, content processing workflows + +## Best Practices + +### System Design + +1. **Start Simple**: Begin with simpler patterns and evolve +2. **Clear Responsibilities**: Define distinct roles for each agent +3. **Robust Communication**: Design reliable message passing +4. **Error Handling**: Plan for failures and recovery +5. **Monitor Everything**: Implement comprehensive observability + +### Tool Design + +1. **Single Responsibility**: Each tool should have one clear purpose +2. **Input Validation**: Validate all inputs thoroughly +3. **Idempotency**: Design operations to be safely repeatable +4. **Error Recovery**: Provide clear error messages and recovery paths +5. **Documentation**: Include comprehensive usage examples + +### Performance Optimization + +1. **Measure First**: Use the evaluator to identify actual bottlenecks +2. **Optimize Bottlenecks**: Focus on highest-impact improvements +3. **Cache Strategically**: Cache expensive operations and results +4. **Parallel Processing**: Identify opportunities for parallelization +5. **Resource Management**: Monitor and optimize resource usage + +## Sample Files + +The `assets/` directory contains sample files to help you get started: + +- **`sample_system_requirements.json`**: Example system requirements for a research platform +- **`sample_tool_descriptions.json`**: Example tool descriptions for common operations +- **`sample_execution_logs.json`**: Example execution logs from a running system + +The `expected_outputs/` directory shows expected results from processing these samples. + +## References + +See the `references/` directory for detailed documentation: + +- **`agent_architecture_patterns.md`**: Comprehensive catalog of architecture patterns +- **`tool_design_best_practices.md`**: Best practices for tool design and implementation +- **`evaluation_methodology.md`**: Detailed methodology for system evaluation + +## Integration Examples + +### With OpenAI + +```python +import json +import openai + +# Load generated OpenAI schemas +with open('my_tools_openai.json') as f: + schemas = json.load(f) + +# Use with OpenAI function calling +response = openai.ChatCompletion.create( + model="gpt-4", + messages=[{"role": "user", "content": "Search for AI news"}], + functions=schemas['functions'] +) +``` + +### With Anthropic Claude + +```python +import json +import anthropic + +# Load generated Anthropic schemas +with open('my_tools_anthropic.json') as f: + schemas = json.load(f) + +# Use with Anthropic tool use +client = anthropic.Anthropic() +response = client.messages.create( + model="claude-3-opus-20240229", + messages=[{"role": "user", "content": "Search for AI news"}], + tools=schemas['tools'] +) +``` + +## Troubleshooting + +### Common Issues + +**"No valid architecture pattern found"** +- Check that team_size is reasonable (1-50) +- Ensure tasks list is not empty +- Verify performance_requirements are valid + +**"Tool schema validation failed"** +- Check that all required fields are present +- Ensure parameter types are valid +- Verify enum values are provided as arrays + +**"Insufficient execution logs"** +- Ensure logs contain required fields (task_id, agent_id, status) +- Check that timestamps are in ISO 8601 format +- Verify token usage fields are numeric + +### Performance Tips + +1. **Large Systems**: For systems with >20 agents, consider breaking into subsystems +2. **Complex Tools**: Tools with >10 parameters may need simplification +3. **Log Volume**: For >1000 log entries, consider sampling for faster analysis + +## Contributing + +This skill is part of the claude-skills repository. To contribute: + +1. Fork the repository +2. Create a feature branch +3. Make your changes +4. Add tests and documentation +5. Submit a pull request + +## License + +This project is licensed under the MIT License - see the main repository for details. + +## Support + +For issues and questions: +- Check the troubleshooting section above +- Review the reference documentation in `references/` +- Create an issue in the claude-skills repository \ No newline at end of file diff --git a/skills/agent-designer/SKILL.md b/skills/agent-designer/SKILL.md new file mode 100644 index 00000000..c9dbd588 --- /dev/null +++ b/skills/agent-designer/SKILL.md @@ -0,0 +1,279 @@ +--- +name: "agent-designer" +description: "Agent Designer - Multi-Agent System Architecture" +--- + +# Agent Designer - Multi-Agent System Architecture + +**Tier:** POWERFUL +**Category:** Engineering +**Tags:** AI agents, architecture, system design, orchestration, multi-agent systems + +## Overview + +Agent Designer is a comprehensive toolkit for designing, architecting, and evaluating multi-agent systems. It provides structured approaches to agent architecture patterns, tool design principles, communication strategies, and performance evaluation frameworks for building robust, scalable AI agent systems. + +## Core Capabilities + +### 1. Agent Architecture Patterns + +#### Single Agent Pattern +- **Use Case:** Simple, focused tasks with clear boundaries +- **Pros:** Minimal complexity, easy debugging, predictable behavior +- **Cons:** Limited scalability, single point of failure +- **Implementation:** Direct user-agent interaction with comprehensive tool access + +#### Supervisor Pattern +- **Use Case:** Hierarchical task decomposition with centralized control +- **Architecture:** One supervisor agent coordinating multiple specialist agents +- **Pros:** Clear command structure, centralized decision making +- **Cons:** Supervisor bottleneck, complex coordination logic +- **Implementation:** Supervisor receives tasks, delegates to specialists, aggregates results + +#### Swarm Pattern +- **Use Case:** Distributed problem solving with peer-to-peer collaboration +- **Architecture:** Multiple autonomous agents with shared objectives +- **Pros:** High parallelism, fault tolerance, emergent intelligence +- **Cons:** Complex coordination, potential conflicts, harder to predict +- **Implementation:** Agent discovery, consensus mechanisms, distributed task allocation + +#### Hierarchical Pattern +- **Use Case:** Complex systems with multiple organizational layers +- **Architecture:** Tree structure with managers and workers at different levels +- **Pros:** Natural organizational mapping, clear responsibilities +- **Cons:** Communication overhead, potential bottlenecks at each level +- **Implementation:** Multi-level delegation with feedback loops + +#### Pipeline Pattern +- **Use Case:** Sequential processing with specialized stages +- **Architecture:** Agents arranged in processing pipeline +- **Pros:** Clear data flow, specialized optimization per stage +- **Cons:** Sequential bottlenecks, rigid processing order +- **Implementation:** Message queues between stages, state handoffs + +### 2. Agent Role Definition + +#### Role Specification Framework +- **Identity:** Name, purpose statement, core competencies +- **Responsibilities:** Primary tasks, decision boundaries, success criteria +- **Capabilities:** Required tools, knowledge domains, processing limits +- **Interfaces:** Input/output formats, communication protocols +- **Constraints:** Security boundaries, resource limits, operational guidelines + +#### Common Agent Archetypes + +**Coordinator Agent** +- Orchestrates multi-agent workflows +- Makes high-level decisions and resource allocation +- Monitors system health and performance +- Handles escalations and conflict resolution + +**Specialist Agent** +- Deep expertise in specific domain (code, data, research) +- Optimized tools and knowledge for specialized tasks +- High-quality output within narrow scope +- Clear handoff protocols for out-of-scope requests + +**Interface Agent** +- Handles external interactions (users, APIs, systems) +- Protocol translation and format conversion +- Authentication and authorization management +- User experience optimization + +**Monitor Agent** +- System health monitoring and alerting +- Performance metrics collection and analysis +- Anomaly detection and reporting +- Compliance and audit trail maintenance + +### 3. Tool Design Principles + +#### Schema Design +- **Input Validation:** Strong typing, required vs optional parameters +- **Output Consistency:** Standardized response formats, error handling +- **Documentation:** Clear descriptions, usage examples, edge cases +- **Versioning:** Backward compatibility, migration paths + +#### Error Handling Patterns +- **Graceful Degradation:** Partial functionality when dependencies fail +- **Retry Logic:** Exponential backoff, circuit breakers, max attempts +- **Error Propagation:** Structured error responses, error classification +- **Recovery Strategies:** Fallback methods, alternative approaches + +#### Idempotency Requirements +- **Safe Operations:** Read operations with no side effects +- **Idempotent Writes:** Same operation can be safely repeated +- **State Management:** Version tracking, conflict resolution +- **Atomicity:** All-or-nothing operation completion + +### 4. Communication Patterns + +#### Message Passing +- **Asynchronous Messaging:** Decoupled agents, message queues +- **Message Format:** Structured payloads with metadata +- **Delivery Guarantees:** At-least-once, exactly-once semantics +- **Routing:** Direct messaging, publish-subscribe, broadcast + +#### Shared State +- **State Stores:** Centralized data repositories +- **Consistency Models:** Strong, eventual, weak consistency +- **Access Patterns:** Read-heavy, write-heavy, mixed workloads +- **Conflict Resolution:** Last-writer-wins, merge strategies + +#### Event-Driven Architecture +- **Event Sourcing:** Immutable event logs, state reconstruction +- **Event Types:** Domain events, system events, integration events +- **Event Processing:** Real-time, batch, stream processing +- **Event Schema:** Versioned event formats, backward compatibility + +### 5. Guardrails and Safety + +#### Input Validation +- **Schema Enforcement:** Required fields, type checking, format validation +- **Content Filtering:** Harmful content detection, PII scrubbing +- **Rate Limiting:** Request throttling, resource quotas +- **Authentication:** Identity verification, authorization checks + +#### Output Filtering +- **Content Moderation:** Harmful content removal, quality checks +- **Consistency Validation:** Logic checks, constraint verification +- **Formatting:** Standardized output formats, clean presentation +- **Audit Logging:** Decision trails, compliance records + +#### Human-in-the-Loop +- **Approval Workflows:** Critical decision checkpoints +- **Escalation Triggers:** Confidence thresholds, risk assessment +- **Override Mechanisms:** Human judgment precedence +- **Feedback Loops:** Human corrections improve system behavior + +### 6. Evaluation Frameworks + +#### Task Completion Metrics +- **Success Rate:** Percentage of tasks completed successfully +- **Partial Completion:** Progress measurement for complex tasks +- **Task Classification:** Success criteria by task type +- **Failure Analysis:** Root cause identification and categorization + +#### Quality Assessment +- **Output Quality:** Accuracy, relevance, completeness measures +- **Consistency:** Response variability across similar inputs +- **Coherence:** Logical flow and internal consistency +- **User Satisfaction:** Feedback scores, usage patterns + +#### Cost Analysis +- **Token Usage:** Input/output token consumption per task +- **API Costs:** External service usage and charges +- **Compute Resources:** CPU, memory, storage utilization +- **Time-to-Value:** Cost per successful task completion + +#### Latency Distribution +- **Response Time:** End-to-end task completion time +- **Processing Stages:** Bottleneck identification per stage +- **Queue Times:** Wait times in processing pipelines +- **Resource Contention:** Impact of concurrent operations + +### 7. Orchestration Strategies + +#### Centralized Orchestration +- **Workflow Engine:** Central coordinator manages all agents +- **State Management:** Centralized workflow state tracking +- **Decision Logic:** Complex routing and branching rules +- **Monitoring:** Comprehensive visibility into all operations + +#### Decentralized Orchestration +- **Peer-to-Peer:** Agents coordinate directly with each other +- **Service Discovery:** Dynamic agent registration and lookup +- **Consensus Protocols:** Distributed decision making +- **Fault Tolerance:** No single point of failure + +#### Hybrid Approaches +- **Domain Boundaries:** Centralized within domains, federated across +- **Hierarchical Coordination:** Multiple orchestration levels +- **Context-Dependent:** Strategy selection based on task type +- **Load Balancing:** Distribute coordination responsibility + +### 8. Memory Patterns + +#### Short-Term Memory +- **Context Windows:** Working memory for current tasks +- **Session State:** Temporary data for ongoing interactions +- **Cache Management:** Performance optimization strategies +- **Memory Pressure:** Handling capacity constraints + +#### Long-Term Memory +- **Persistent Storage:** Durable data across sessions +- **Knowledge Base:** Accumulated domain knowledge +- **Experience Replay:** Learning from past interactions +- **Memory Consolidation:** Transferring from short to long-term + +#### Shared Memory +- **Collaborative Knowledge:** Shared learning across agents +- **Synchronization:** Consistency maintenance strategies +- **Access Control:** Permission-based memory access +- **Memory Partitioning:** Isolation between agent groups + +### 9. Scaling Considerations + +#### Horizontal Scaling +- **Agent Replication:** Multiple instances of same agent type +- **Load Distribution:** Request routing across agent instances +- **Resource Pooling:** Shared compute and storage resources +- **Geographic Distribution:** Multi-region deployments + +#### Vertical Scaling +- **Capability Enhancement:** More powerful individual agents +- **Tool Expansion:** Broader tool access per agent +- **Context Expansion:** Larger working memory capacity +- **Processing Power:** Higher throughput per agent + +#### Performance Optimization +- **Caching Strategies:** Response caching, tool result caching +- **Parallel Processing:** Concurrent task execution +- **Resource Optimization:** Efficient resource utilization +- **Bottleneck Elimination:** Systematic performance tuning + +### 10. Failure Handling + +#### Retry Mechanisms +- **Exponential Backoff:** Increasing delays between retries +- **Jitter:** Random delay variation to prevent thundering herd +- **Maximum Attempts:** Bounded retry behavior +- **Retry Conditions:** Transient vs permanent failure classification + +#### Fallback Strategies +- **Graceful Degradation:** Reduced functionality when systems fail +- **Alternative Approaches:** Different methods for same goals +- **Default Responses:** Safe fallback behaviors +- **User Communication:** Clear failure messaging + +#### Circuit Breakers +- **Failure Detection:** Monitoring failure rates and response times +- **State Management:** Open, closed, half-open circuit states +- **Recovery Testing:** Gradual return to normal operation +- **Cascading Failure Prevention:** Protecting upstream systems + +## Implementation Guidelines + +### Architecture Decision Process +1. **Requirements Analysis:** Understand system goals, constraints, scale +2. **Pattern Selection:** Choose appropriate architecture pattern +3. **Agent Design:** Define roles, responsibilities, interfaces +4. **Tool Architecture:** Design tool schemas and error handling +5. **Communication Design:** Select message patterns and protocols +6. **Safety Implementation:** Build guardrails and validation +7. **Evaluation Planning:** Define success metrics and monitoring +8. **Deployment Strategy:** Plan scaling and failure handling + +### Quality Assurance +- **Testing Strategy:** Unit, integration, and system testing approaches +- **Monitoring:** Real-time system health and performance tracking +- **Documentation:** Architecture documentation and runbooks +- **Security Review:** Threat modeling and security assessments + +### Continuous Improvement +- **Performance Monitoring:** Ongoing system performance analysis +- **User Feedback:** Incorporating user experience improvements +- **A/B Testing:** Controlled experiments for system improvements +- **Knowledge Base Updates:** Continuous learning and adaptation + +This skill provides the foundation for designing robust, scalable multi-agent systems that can handle complex tasks while maintaining safety, reliability, and performance at scale. \ No newline at end of file diff --git a/skills/agent-designer/_meta.json b/skills/agent-designer/_meta.json new file mode 100644 index 00000000..6eff9633 --- /dev/null +++ b/skills/agent-designer/_meta.json @@ -0,0 +1,17 @@ +{ + "owner": "alirezarezvani", + "slug": "agent-designer", + "displayName": "Agent Designer", + "latest": { + "version": "2.1.1", + "publishedAt": 1773120983263, + "commit": "https://github.com/openclaw/skills/commit/7efd6bf8ba027c6ce5beb82c06fda1ba0dea136e" + }, + "history": [ + { + "version": "1.0.0", + "publishedAt": 1771270589254, + "commit": "https://github.com/openclaw/skills/commit/f66cebd15374a316825c2cff3b61bbeafa4dbd8a" + } + ] +} diff --git a/skills/agent-designer/agent_evaluator.py b/skills/agent-designer/agent_evaluator.py new file mode 100644 index 00000000..709171c9 --- /dev/null +++ b/skills/agent-designer/agent_evaluator.py @@ -0,0 +1,1223 @@ +#!/usr/bin/env python3 +""" +Agent Evaluator - Multi-Agent System Performance Analysis + +Takes agent execution logs (task, actions taken, results, time, tokens used) +and evaluates performance: task success rate, average cost per task, latency +distribution, error patterns, tool usage efficiency, identifies bottlenecks +and improvement opportunities. + +Input: execution logs JSON +Output: performance report + bottleneck analysis + optimization recommendations +""" + +import json +import argparse +import sys +import statistics +from typing import Dict, List, Any, Optional, Tuple +from dataclasses import dataclass, asdict +from collections import defaultdict, Counter +from datetime import datetime, timedelta +import re + + +@dataclass +class ExecutionLog: + """Single execution log entry""" + task_id: str + agent_id: str + task_type: str + task_description: str + start_time: str + end_time: str + duration_ms: int + status: str # success, failure, partial, timeout + actions: List[Dict[str, Any]] + results: Dict[str, Any] + tokens_used: Dict[str, int] # input_tokens, output_tokens, total_tokens + cost_usd: float + error_details: Optional[Dict[str, Any]] + tools_used: List[str] + retry_count: int + metadata: Dict[str, Any] + + +@dataclass +class PerformanceMetrics: + """Performance metrics for an agent or system""" + total_tasks: int + successful_tasks: int + failed_tasks: int + partial_tasks: int + timeout_tasks: int + success_rate: float + failure_rate: float + average_duration_ms: float + median_duration_ms: float + percentile_95_duration_ms: float + min_duration_ms: int + max_duration_ms: int + total_tokens_used: int + average_tokens_per_task: float + total_cost_usd: float + average_cost_per_task: float + cost_per_token: float + throughput_tasks_per_hour: float + error_rate: float + retry_rate: float + + +@dataclass +class ErrorAnalysis: + """Error pattern analysis""" + error_type: str + count: int + percentage: float + affected_agents: List[str] + affected_task_types: List[str] + common_patterns: List[str] + suggested_fixes: List[str] + impact_level: str # high, medium, low + + +@dataclass +class BottleneckAnalysis: + """System bottleneck analysis""" + bottleneck_type: str # agent, tool, communication, resource + location: str + severity: str # critical, high, medium, low + description: str + impact_on_performance: Dict[str, float] + affected_workflows: List[str] + optimization_suggestions: List[str] + estimated_improvement: Dict[str, float] + + +@dataclass +class OptimizationRecommendation: + """Performance optimization recommendation""" + category: str # performance, cost, reliability, scalability + priority: str # high, medium, low + title: str + description: str + implementation_effort: str # low, medium, high + expected_impact: Dict[str, Any] + estimated_cost_savings: Optional[float] + estimated_performance_gain: Optional[float] + implementation_steps: List[str] + risks: List[str] + prerequisites: List[str] + + +@dataclass +class EvaluationReport: + """Complete evaluation report""" + summary: Dict[str, Any] + system_metrics: PerformanceMetrics + agent_metrics: Dict[str, PerformanceMetrics] + task_type_metrics: Dict[str, PerformanceMetrics] + tool_usage_analysis: Dict[str, Any] + error_analysis: List[ErrorAnalysis] + bottleneck_analysis: List[BottleneckAnalysis] + optimization_recommendations: List[OptimizationRecommendation] + trends_analysis: Dict[str, Any] + cost_breakdown: Dict[str, Any] + sla_compliance: Dict[str, Any] + metadata: Dict[str, Any] + + +class AgentEvaluator: + """Evaluate multi-agent system performance from execution logs""" + + def __init__(self): + self.error_patterns = self._define_error_patterns() + self.performance_thresholds = self._define_performance_thresholds() + self.cost_benchmarks = self._define_cost_benchmarks() + + def _define_error_patterns(self) -> Dict[str, Dict[str, Any]]: + """Define common error patterns and their classifications""" + return { + "timeout": { + "patterns": [r"timeout", r"timed out", r"deadline exceeded"], + "category": "performance", + "severity": "high", + "common_fixes": [ + "Increase timeout values", + "Optimize slow operations", + "Add retry logic with exponential backoff", + "Parallelize independent operations" + ] + }, + "rate_limit": { + "patterns": [r"rate limit", r"too many requests", r"quota exceeded"], + "category": "resource", + "severity": "medium", + "common_fixes": [ + "Implement request throttling", + "Add circuit breaker pattern", + "Use request queuing", + "Negotiate higher limits" + ] + }, + "authentication": { + "patterns": [r"unauthorized", r"authentication failed", r"invalid credentials"], + "category": "security", + "severity": "high", + "common_fixes": [ + "Check credential rotation", + "Implement token refresh logic", + "Add authentication retry", + "Verify permission scopes" + ] + }, + "network": { + "patterns": [r"connection refused", r"network error", r"dns resolution"], + "category": "infrastructure", + "severity": "high", + "common_fixes": [ + "Add network retry logic", + "Implement fallback endpoints", + "Use connection pooling", + "Add health checks" + ] + }, + "validation": { + "patterns": [r"validation error", r"invalid input", r"schema violation"], + "category": "data", + "severity": "medium", + "common_fixes": [ + "Strengthen input validation", + "Add data sanitization", + "Improve error messages", + "Add input examples" + ] + }, + "resource": { + "patterns": [r"out of memory", r"disk full", r"cpu overload"], + "category": "resource", + "severity": "critical", + "common_fixes": [ + "Scale up resources", + "Optimize memory usage", + "Add resource monitoring", + "Implement graceful degradation" + ] + } + } + + def _define_performance_thresholds(self) -> Dict[str, Any]: + """Define performance thresholds for different metrics""" + return { + "success_rate": {"excellent": 0.98, "good": 0.95, "acceptable": 0.90, "poor": 0.80}, + "average_duration": {"excellent": 1000, "good": 3000, "acceptable": 10000, "poor": 30000}, + "error_rate": {"excellent": 0.01, "good": 0.03, "acceptable": 0.05, "poor": 0.10}, + "retry_rate": {"excellent": 0.05, "good": 0.10, "acceptable": 0.20, "poor": 0.40}, + "cost_per_task": {"excellent": 0.01, "good": 0.05, "acceptable": 0.10, "poor": 0.25}, + "throughput": {"excellent": 100, "good": 50, "acceptable": 20, "poor": 5} # tasks per hour + } + + def _define_cost_benchmarks(self) -> Dict[str, Any]: + """Define cost benchmarks for different operations""" + return { + "token_costs": { + "gpt-4": {"input": 0.00003, "output": 0.00006}, + "gpt-3.5-turbo": {"input": 0.000002, "output": 0.000002}, + "claude-3": {"input": 0.000015, "output": 0.000075} + }, + "operation_costs": { + "simple_task": 0.005, + "complex_task": 0.050, + "research_task": 0.020, + "analysis_task": 0.030, + "generation_task": 0.015 + } + } + + def parse_execution_logs(self, logs_data: List[Dict[str, Any]]) -> List[ExecutionLog]: + """Parse raw execution logs into structured format""" + logs = [] + + for log_entry in logs_data: + try: + log = ExecutionLog( + task_id=log_entry.get("task_id", ""), + agent_id=log_entry.get("agent_id", ""), + task_type=log_entry.get("task_type", "unknown"), + task_description=log_entry.get("task_description", ""), + start_time=log_entry.get("start_time", ""), + end_time=log_entry.get("end_time", ""), + duration_ms=log_entry.get("duration_ms", 0), + status=log_entry.get("status", "unknown"), + actions=log_entry.get("actions", []), + results=log_entry.get("results", {}), + tokens_used=log_entry.get("tokens_used", {"total_tokens": 0}), + cost_usd=log_entry.get("cost_usd", 0.0), + error_details=log_entry.get("error_details"), + tools_used=log_entry.get("tools_used", []), + retry_count=log_entry.get("retry_count", 0), + metadata=log_entry.get("metadata", {}) + ) + logs.append(log) + except Exception as e: + print(f"Warning: Failed to parse log entry: {e}", file=sys.stderr) + continue + + return logs + + def calculate_performance_metrics(self, logs: List[ExecutionLog]) -> PerformanceMetrics: + """Calculate performance metrics from execution logs""" + if not logs: + return PerformanceMetrics( + total_tasks=0, successful_tasks=0, failed_tasks=0, partial_tasks=0, + timeout_tasks=0, success_rate=0.0, failure_rate=0.0, + average_duration_ms=0.0, median_duration_ms=0.0, percentile_95_duration_ms=0.0, + min_duration_ms=0, max_duration_ms=0, total_tokens_used=0, + average_tokens_per_task=0.0, total_cost_usd=0.0, average_cost_per_task=0.0, + cost_per_token=0.0, throughput_tasks_per_hour=0.0, error_rate=0.0, retry_rate=0.0 + ) + + total_tasks = len(logs) + successful_tasks = sum(1 for log in logs if log.status == "success") + failed_tasks = sum(1 for log in logs if log.status == "failure") + partial_tasks = sum(1 for log in logs if log.status == "partial") + timeout_tasks = sum(1 for log in logs if log.status == "timeout") + + success_rate = successful_tasks / total_tasks if total_tasks > 0 else 0.0 + failure_rate = (failed_tasks + timeout_tasks) / total_tasks if total_tasks > 0 else 0.0 + + durations = [log.duration_ms for log in logs if log.duration_ms > 0] + if durations: + average_duration_ms = statistics.mean(durations) + median_duration_ms = statistics.median(durations) + percentile_95_duration_ms = self._percentile(durations, 95) + min_duration_ms = min(durations) + max_duration_ms = max(durations) + else: + average_duration_ms = median_duration_ms = percentile_95_duration_ms = 0.0 + min_duration_ms = max_duration_ms = 0 + + total_tokens = sum(log.tokens_used.get("total_tokens", 0) for log in logs) + average_tokens_per_task = total_tokens / total_tasks if total_tasks > 0 else 0.0 + + total_cost = sum(log.cost_usd for log in logs) + average_cost_per_task = total_cost / total_tasks if total_tasks > 0 else 0.0 + cost_per_token = total_cost / total_tokens if total_tokens > 0 else 0.0 + + # Calculate throughput (tasks per hour) + if logs and len(logs) > 1: + start_time = min(log.start_time for log in logs if log.start_time) + end_time = max(log.end_time for log in logs if log.end_time) + if start_time and end_time: + try: + start_dt = datetime.fromisoformat(start_time.replace("Z", "+00:00")) + end_dt = datetime.fromisoformat(end_time.replace("Z", "+00:00")) + time_diff_hours = (end_dt - start_dt).total_seconds() / 3600 + throughput_tasks_per_hour = total_tasks / time_diff_hours if time_diff_hours > 0 else 0.0 + except: + throughput_tasks_per_hour = 0.0 + else: + throughput_tasks_per_hour = 0.0 + else: + throughput_tasks_per_hour = 0.0 + + error_rate = sum(1 for log in logs if log.error_details) / total_tasks if total_tasks > 0 else 0.0 + retry_rate = sum(1 for log in logs if log.retry_count > 0) / total_tasks if total_tasks > 0 else 0.0 + + return PerformanceMetrics( + total_tasks=total_tasks, + successful_tasks=successful_tasks, + failed_tasks=failed_tasks, + partial_tasks=partial_tasks, + timeout_tasks=timeout_tasks, + success_rate=success_rate, + failure_rate=failure_rate, + average_duration_ms=average_duration_ms, + median_duration_ms=median_duration_ms, + percentile_95_duration_ms=percentile_95_duration_ms, + min_duration_ms=min_duration_ms, + max_duration_ms=max_duration_ms, + total_tokens_used=total_tokens, + average_tokens_per_task=average_tokens_per_task, + total_cost_usd=total_cost, + average_cost_per_task=average_cost_per_task, + cost_per_token=cost_per_token, + throughput_tasks_per_hour=throughput_tasks_per_hour, + error_rate=error_rate, + retry_rate=retry_rate + ) + + def _percentile(self, data: List[float], percentile: int) -> float: + """Calculate percentile value from data""" + if not data: + return 0.0 + sorted_data = sorted(data) + index = (percentile / 100) * (len(sorted_data) - 1) + if index.is_integer(): + return sorted_data[int(index)] + else: + lower_index = int(index) + upper_index = lower_index + 1 + weight = index - lower_index + return sorted_data[lower_index] * (1 - weight) + sorted_data[upper_index] * weight + + def analyze_errors(self, logs: List[ExecutionLog]) -> List[ErrorAnalysis]: + """Analyze error patterns in execution logs""" + error_analyses = [] + + # Collect all errors + errors = [] + for log in logs: + if log.error_details: + errors.append({ + "error": log.error_details, + "agent_id": log.agent_id, + "task_type": log.task_type, + "task_id": log.task_id + }) + + if not errors: + return error_analyses + + # Group errors by pattern + error_groups = defaultdict(list) + unclassified_errors = [] + + for error in errors: + error_message = str(error.get("error", {})).lower() + classified = False + + for pattern_name, pattern_info in self.error_patterns.items(): + for pattern in pattern_info["patterns"]: + if re.search(pattern, error_message): + error_groups[pattern_name].append(error) + classified = True + break + if classified: + break + + if not classified: + unclassified_errors.append(error) + + # Analyze each error group + total_errors = len(errors) + + for error_type, error_list in error_groups.items(): + count = len(error_list) + percentage = (count / total_errors) * 100 if total_errors > 0 else 0.0 + + affected_agents = list(set(error["agent_id"] for error in error_list)) + affected_task_types = list(set(error["task_type"] for error in error_list)) + + # Extract common patterns from error messages + common_patterns = self._extract_common_patterns([str(e["error"]) for e in error_list]) + + # Get suggested fixes + pattern_info = self.error_patterns.get(error_type, {}) + suggested_fixes = pattern_info.get("common_fixes", []) + + # Determine impact level + if percentage > 20 or pattern_info.get("severity") == "critical": + impact_level = "high" + elif percentage > 10 or pattern_info.get("severity") == "high": + impact_level = "medium" + else: + impact_level = "low" + + error_analysis = ErrorAnalysis( + error_type=error_type, + count=count, + percentage=percentage, + affected_agents=affected_agents, + affected_task_types=affected_task_types, + common_patterns=common_patterns, + suggested_fixes=suggested_fixes, + impact_level=impact_level + ) + + error_analyses.append(error_analysis) + + # Handle unclassified errors + if unclassified_errors: + count = len(unclassified_errors) + percentage = (count / total_errors) * 100 + + error_analysis = ErrorAnalysis( + error_type="unclassified", + count=count, + percentage=percentage, + affected_agents=list(set(error["agent_id"] for error in unclassified_errors)), + affected_task_types=list(set(error["task_type"] for error in unclassified_errors)), + common_patterns=self._extract_common_patterns([str(e["error"]) for e in unclassified_errors]), + suggested_fixes=["Review and classify error patterns", "Add specific error handling"], + impact_level="medium" if percentage > 10 else "low" + ) + + error_analyses.append(error_analysis) + + # Sort by impact and count + error_analyses.sort(key=lambda x: (x.impact_level == "high", x.count), reverse=True) + + return error_analyses + + def _extract_common_patterns(self, error_messages: List[str]) -> List[str]: + """Extract common patterns from error messages""" + if not error_messages: + return [] + + # Simple pattern extraction - find common phrases + word_counts = Counter() + for message in error_messages: + words = re.findall(r'\w+', message.lower()) + for word in words: + if len(word) > 3: # Ignore short words + word_counts[word] += 1 + + # Return most common words/patterns + common_patterns = [word for word, count in word_counts.most_common(5) + if count > 1] + + return common_patterns + + def identify_bottlenecks(self, logs: List[ExecutionLog], + agent_metrics: Dict[str, PerformanceMetrics]) -> List[BottleneckAnalysis]: + """Identify system bottlenecks""" + bottlenecks = [] + + # Agent performance bottlenecks + for agent_id, metrics in agent_metrics.items(): + if metrics.success_rate < 0.8: + severity = "critical" if metrics.success_rate < 0.5 else "high" + bottlenecks.append(BottleneckAnalysis( + bottleneck_type="agent", + location=agent_id, + severity=severity, + description=f"Agent {agent_id} has low success rate ({metrics.success_rate:.1%})", + impact_on_performance={ + "success_rate_impact": (0.95 - metrics.success_rate) * 100, + "cost_impact": metrics.average_cost_per_task * metrics.failed_tasks + }, + affected_workflows=self._get_agent_workflows(agent_id, logs), + optimization_suggestions=[ + "Review and improve agent logic", + "Add better error handling", + "Optimize tool usage", + "Consider agent specialization" + ], + estimated_improvement={ + "success_rate_gain": min(0.15, 0.95 - metrics.success_rate), + "cost_reduction": metrics.average_cost_per_task * 0.2 + } + )) + + if metrics.average_duration_ms > 30000: # 30 seconds + severity = "high" if metrics.average_duration_ms > 60000 else "medium" + bottlenecks.append(BottleneckAnalysis( + bottleneck_type="agent", + location=agent_id, + severity=severity, + description=f"Agent {agent_id} has high latency ({metrics.average_duration_ms/1000:.1f}s avg)", + impact_on_performance={ + "latency_impact": metrics.average_duration_ms - 10000, + "throughput_impact": max(0, 50 - metrics.total_tasks) + }, + affected_workflows=self._get_agent_workflows(agent_id, logs), + optimization_suggestions=[ + "Profile and optimize slow operations", + "Implement caching strategies", + "Parallelize independent tasks", + "Optimize API calls" + ], + estimated_improvement={ + "latency_reduction": min(0.5, (metrics.average_duration_ms - 10000) / metrics.average_duration_ms), + "throughput_gain": 1.3 + } + )) + + # Tool usage bottlenecks + tool_usage = self._analyze_tool_usage(logs) + for tool, usage_stats in tool_usage.items(): + if usage_stats.get("error_rate", 0) > 0.2: + bottlenecks.append(BottleneckAnalysis( + bottleneck_type="tool", + location=tool, + severity="high" if usage_stats["error_rate"] > 0.4 else "medium", + description=f"Tool {tool} has high error rate ({usage_stats['error_rate']:.1%})", + impact_on_performance={ + "reliability_impact": usage_stats["error_rate"] * usage_stats["usage_count"], + "retry_overhead": usage_stats.get("retry_count", 0) * 1000 # ms + }, + affected_workflows=usage_stats.get("affected_workflows", []), + optimization_suggestions=[ + "Review tool implementation", + "Add better error handling for tool", + "Implement tool fallbacks", + "Consider alternative tools" + ], + estimated_improvement={ + "error_reduction": usage_stats["error_rate"] * 0.7, + "performance_gain": 1.2 + } + )) + + # Communication bottlenecks + communication_analysis = self._analyze_communication_patterns(logs) + if communication_analysis.get("high_latency_communications", 0) > 5: + bottlenecks.append(BottleneckAnalysis( + bottleneck_type="communication", + location="inter_agent_communication", + severity="medium", + description="High latency in inter-agent communications detected", + impact_on_performance={ + "communication_overhead": communication_analysis.get("avg_communication_latency", 0), + "coordination_efficiency": 0.8 # Assumed impact + }, + affected_workflows=communication_analysis.get("affected_workflows", []), + optimization_suggestions=[ + "Optimize message serialization", + "Implement message batching", + "Add communication caching", + "Consider direct communication patterns" + ], + estimated_improvement={ + "communication_latency_reduction": 0.4, + "overall_efficiency_gain": 1.15 + } + )) + + # Resource bottlenecks + resource_analysis = self._analyze_resource_usage(logs) + if resource_analysis.get("high_token_usage_tasks", 0) > 10: + bottlenecks.append(BottleneckAnalysis( + bottleneck_type="resource", + location="token_usage", + severity="medium", + description="High token usage detected in multiple tasks", + impact_on_performance={ + "cost_impact": resource_analysis.get("excess_token_cost", 0), + "latency_impact": resource_analysis.get("token_processing_overhead", 0) + }, + affected_workflows=resource_analysis.get("high_usage_workflows", []), + optimization_suggestions=[ + "Optimize prompt engineering", + "Implement response caching", + "Use more efficient models for simple tasks", + "Add token usage monitoring" + ], + estimated_improvement={ + "cost_reduction": 0.3, + "efficiency_gain": 1.1 + } + )) + + # Sort bottlenecks by severity and impact + severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3} + bottlenecks.sort(key=lambda x: (severity_order[x.severity], + -sum(x.impact_on_performance.values()))) + + return bottlenecks + + def _get_agent_workflows(self, agent_id: str, logs: List[ExecutionLog]) -> List[str]: + """Get workflows affected by a specific agent""" + workflows = set() + for log in logs: + if log.agent_id == agent_id: + workflows.add(log.task_type) + return list(workflows) + + def _analyze_tool_usage(self, logs: List[ExecutionLog]) -> Dict[str, Dict[str, Any]]: + """Analyze tool usage patterns""" + tool_stats = defaultdict(lambda: { + "usage_count": 0, + "error_count": 0, + "total_duration": 0, + "affected_workflows": set(), + "retry_count": 0 + }) + + for log in logs: + for tool in log.tools_used: + stats = tool_stats[tool] + stats["usage_count"] += 1 + stats["total_duration"] += log.duration_ms + stats["affected_workflows"].add(log.task_type) + + if log.error_details: + stats["error_count"] += 1 + if log.retry_count > 0: + stats["retry_count"] += log.retry_count + + # Calculate derived metrics + result = {} + for tool, stats in tool_stats.items(): + result[tool] = { + "usage_count": stats["usage_count"], + "error_rate": stats["error_count"] / stats["usage_count"] if stats["usage_count"] > 0 else 0, + "avg_duration": stats["total_duration"] / stats["usage_count"] if stats["usage_count"] > 0 else 0, + "affected_workflows": list(stats["affected_workflows"]), + "retry_count": stats["retry_count"] + } + + return result + + def _analyze_communication_patterns(self, logs: List[ExecutionLog]) -> Dict[str, Any]: + """Analyze communication patterns between agents""" + # This is a simplified analysis - in a real system, you'd have more detailed communication logs + communication_actions = [] + for log in logs: + for action in log.actions: + if action.get("type") in ["message", "delegate", "coordinate", "respond"]: + communication_actions.append({ + "duration": action.get("duration_ms", 0), + "success": action.get("success", True), + "workflow": log.task_type + }) + + if not communication_actions: + return {} + + avg_latency = sum(action["duration"] for action in communication_actions) / len(communication_actions) + high_latency_count = sum(1 for action in communication_actions if action["duration"] > 5000) + + return { + "total_communications": len(communication_actions), + "avg_communication_latency": avg_latency, + "high_latency_communications": high_latency_count, + "affected_workflows": list(set(action["workflow"] for action in communication_actions)) + } + + def _analyze_resource_usage(self, logs: List[ExecutionLog]) -> Dict[str, Any]: + """Analyze resource usage patterns""" + token_usage = [log.tokens_used.get("total_tokens", 0) for log in logs] + + if not token_usage: + return {} + + avg_tokens = sum(token_usage) / len(token_usage) + high_usage_threshold = avg_tokens * 2 + high_usage_tasks = sum(1 for tokens in token_usage if tokens > high_usage_threshold) + + # Estimate excess cost + excess_tokens = sum(max(0, tokens - avg_tokens) for tokens in token_usage) + excess_cost = excess_tokens * 0.00002 # Rough estimate + + return { + "avg_token_usage": avg_tokens, + "high_token_usage_tasks": high_usage_tasks, + "excess_token_cost": excess_cost, + "token_processing_overhead": high_usage_tasks * 500, # Estimated overhead in ms + "high_usage_workflows": [log.task_type for log in logs + if log.tokens_used.get("total_tokens", 0) > high_usage_threshold] + } + + def generate_optimization_recommendations(self, + system_metrics: PerformanceMetrics, + error_analyses: List[ErrorAnalysis], + bottlenecks: List[BottleneckAnalysis]) -> List[OptimizationRecommendation]: + """Generate optimization recommendations based on analysis""" + recommendations = [] + + # Performance optimization recommendations + if system_metrics.success_rate < 0.9: + recommendations.append(OptimizationRecommendation( + category="reliability", + priority="high", + title="Improve System Reliability", + description=f"System success rate is {system_metrics.success_rate:.1%}, below target of 90%", + implementation_effort="medium", + expected_impact={ + "success_rate_improvement": min(0.1, 0.95 - system_metrics.success_rate), + "cost_reduction": system_metrics.average_cost_per_task * 0.15 + }, + estimated_cost_savings=system_metrics.total_cost_usd * 0.1, + estimated_performance_gain=1.2, + implementation_steps=[ + "Identify and fix top error patterns", + "Implement better error handling and retries", + "Add comprehensive monitoring and alerting", + "Implement graceful degradation patterns" + ], + risks=["Temporary increase in complexity", "Potential initial performance overhead"], + prerequisites=["Error analysis completion", "Monitoring infrastructure"] + )) + + # Cost optimization recommendations + if system_metrics.average_cost_per_task > 0.1: + recommendations.append(OptimizationRecommendation( + category="cost", + priority="medium", + title="Optimize Token Usage and Costs", + description=f"Average cost per task (${system_metrics.average_cost_per_task:.3f}) is above optimal range", + implementation_effort="low", + expected_impact={ + "cost_reduction": system_metrics.average_cost_per_task * 0.3, + "efficiency_improvement": 1.15 + }, + estimated_cost_savings=system_metrics.total_cost_usd * 0.3, + estimated_performance_gain=1.05, + implementation_steps=[ + "Implement prompt optimization", + "Add response caching for repeated queries", + "Use smaller models for simple tasks", + "Implement token usage monitoring and alerts" + ], + risks=["Potential quality reduction with smaller models"], + prerequisites=["Token usage analysis", "Caching infrastructure"] + )) + + # Performance optimization recommendations + if system_metrics.average_duration_ms > 10000: + recommendations.append(OptimizationRecommendation( + category="performance", + priority="high", + title="Reduce Task Latency", + description=f"Average task duration ({system_metrics.average_duration_ms/1000:.1f}s) exceeds target", + implementation_effort="high", + expected_impact={ + "latency_reduction": min(0.5, (system_metrics.average_duration_ms - 5000) / system_metrics.average_duration_ms), + "throughput_improvement": 1.5 + }, + estimated_performance_gain=1.4, + implementation_steps=[ + "Profile and optimize slow operations", + "Implement parallel processing where possible", + "Add caching for expensive operations", + "Optimize API calls and reduce round trips" + ], + risks=["Increased system complexity", "Potential resource usage increase"], + prerequisites=["Performance profiling tools", "Caching infrastructure"] + )) + + # Error-based recommendations + high_impact_errors = [ea for ea in error_analyses if ea.impact_level == "high"] + if high_impact_errors: + for error_analysis in high_impact_errors[:3]: # Top 3 high impact errors + recommendations.append(OptimizationRecommendation( + category="reliability", + priority="high", + title=f"Address {error_analysis.error_type.title()} Errors", + description=f"{error_analysis.error_type.title()} errors occur in {error_analysis.percentage:.1f}% of cases", + implementation_effort="medium", + expected_impact={ + "error_reduction": error_analysis.percentage / 100, + "reliability_improvement": 1.1 + }, + estimated_cost_savings=system_metrics.total_cost_usd * (error_analysis.percentage / 100) * 0.5, + implementation_steps=error_analysis.suggested_fixes, + risks=["May require significant code changes"], + prerequisites=["Root cause analysis", "Testing framework"] + )) + + # Bottleneck-based recommendations + critical_bottlenecks = [b for b in bottlenecks if b.severity in ["critical", "high"]] + for bottleneck in critical_bottlenecks[:2]: # Top 2 critical bottlenecks + recommendations.append(OptimizationRecommendation( + category="performance", + priority="high" if bottleneck.severity == "critical" else "medium", + title=f"Address {bottleneck.bottleneck_type.title()} Bottleneck", + description=bottleneck.description, + implementation_effort="medium", + expected_impact=bottleneck.estimated_improvement, + estimated_performance_gain=list(bottleneck.estimated_improvement.values())[0] if bottleneck.estimated_improvement else 1.1, + implementation_steps=bottleneck.optimization_suggestions, + risks=["System downtime during implementation", "Potential cascade effects"], + prerequisites=["Impact assessment", "Rollback plan"] + )) + + # Scalability recommendations + if system_metrics.throughput_tasks_per_hour < 20: + recommendations.append(OptimizationRecommendation( + category="scalability", + priority="medium", + title="Improve System Scalability", + description="Current throughput indicates potential scalability issues", + implementation_effort="high", + expected_impact={ + "throughput_improvement": 2.0, + "scalability_headroom": 5.0 + }, + estimated_performance_gain=2.0, + implementation_steps=[ + "Implement horizontal scaling for agents", + "Add load balancing and resource pooling", + "Optimize resource allocation algorithms", + "Implement auto-scaling policies" + ], + risks=["High implementation complexity", "Increased operational overhead"], + prerequisites=["Infrastructure scaling capability", "Monitoring and metrics"] + )) + + # Sort recommendations by priority and impact + priority_order = {"high": 0, "medium": 1, "low": 2} + recommendations.sort(key=lambda x: ( + priority_order[x.priority], + -x.estimated_performance_gain if x.estimated_performance_gain else 0, + -x.estimated_cost_savings if x.estimated_cost_savings else 0 + )) + + return recommendations + + def generate_report(self, logs: List[ExecutionLog]) -> EvaluationReport: + """Generate complete evaluation report""" + + # Calculate system metrics + system_metrics = self.calculate_performance_metrics(logs) + + # Calculate per-agent metrics + agents = set(log.agent_id for log in logs) + agent_metrics = {} + for agent_id in agents: + agent_logs = [log for log in logs if log.agent_id == agent_id] + agent_metrics[agent_id] = self.calculate_performance_metrics(agent_logs) + + # Calculate per-task-type metrics + task_types = set(log.task_type for log in logs) + task_type_metrics = {} + for task_type in task_types: + task_logs = [log for log in logs if log.task_type == task_type] + task_type_metrics[task_type] = self.calculate_performance_metrics(task_logs) + + # Analyze tool usage + tool_usage_analysis = self._analyze_tool_usage(logs) + + # Analyze errors + error_analysis = self.analyze_errors(logs) + + # Identify bottlenecks + bottleneck_analysis = self.identify_bottlenecks(logs, agent_metrics) + + # Generate optimization recommendations + optimization_recommendations = self.generate_optimization_recommendations( + system_metrics, error_analysis, bottleneck_analysis) + + # Generate trends analysis (simplified) + trends_analysis = self._generate_trends_analysis(logs) + + # Generate cost breakdown + cost_breakdown = self._generate_cost_breakdown(logs, agent_metrics) + + # Check SLA compliance + sla_compliance = self._check_sla_compliance(system_metrics) + + # Create summary + summary = { + "evaluation_period": { + "start_time": min(log.start_time for log in logs if log.start_time) if logs else None, + "end_time": max(log.end_time for log in logs if log.end_time) if logs else None, + "total_duration_hours": system_metrics.total_tasks / system_metrics.throughput_tasks_per_hour if system_metrics.throughput_tasks_per_hour > 0 else 0 + }, + "overall_health": self._assess_overall_health(system_metrics), + "key_findings": self._extract_key_findings(system_metrics, error_analysis, bottleneck_analysis), + "critical_issues": len([b for b in bottleneck_analysis if b.severity == "critical"]), + "improvement_opportunities": len(optimization_recommendations) + } + + # Create metadata + metadata = { + "generated_at": datetime.now().isoformat(), + "evaluator_version": "1.0", + "total_logs_processed": len(logs), + "agents_analyzed": len(agents), + "task_types_analyzed": len(task_types), + "analysis_completeness": "full" + } + + return EvaluationReport( + summary=summary, + system_metrics=system_metrics, + agent_metrics=agent_metrics, + task_type_metrics=task_type_metrics, + tool_usage_analysis=tool_usage_analysis, + error_analysis=error_analysis, + bottleneck_analysis=bottleneck_analysis, + optimization_recommendations=optimization_recommendations, + trends_analysis=trends_analysis, + cost_breakdown=cost_breakdown, + sla_compliance=sla_compliance, + metadata=metadata + ) + + def _generate_trends_analysis(self, logs: List[ExecutionLog]) -> Dict[str, Any]: + """Generate trends analysis (simplified version)""" + # Group logs by time periods (daily) + daily_metrics = defaultdict(list) + + for log in logs: + if log.start_time: + try: + date = log.start_time.split('T')[0] # Extract date part + daily_metrics[date].append(log) + except: + continue + + trends = {} + if len(daily_metrics) > 1: + daily_success_rates = {} + daily_avg_durations = {} + daily_costs = {} + + for date, date_logs in daily_metrics.items(): + if date_logs: + metrics = self.calculate_performance_metrics(date_logs) + daily_success_rates[date] = metrics.success_rate + daily_avg_durations[date] = metrics.average_duration_ms + daily_costs[date] = metrics.total_cost_usd + + trends = { + "daily_success_rates": daily_success_rates, + "daily_avg_durations": daily_avg_durations, + "daily_costs": daily_costs, + "trend_direction": { + "success_rate": "stable", # Simplified + "duration": "stable", + "cost": "stable" + } + } + + return trends + + def _generate_cost_breakdown(self, logs: List[ExecutionLog], + agent_metrics: Dict[str, PerformanceMetrics]) -> Dict[str, Any]: + """Generate cost breakdown analysis""" + total_cost = sum(log.cost_usd for log in logs) + + # Cost by agent + agent_costs = {} + for agent_id, metrics in agent_metrics.items(): + agent_costs[agent_id] = metrics.total_cost_usd + + # Cost by task type + task_type_costs = defaultdict(float) + for log in logs: + task_type_costs[log.task_type] += log.cost_usd + + # Token cost breakdown + total_tokens = sum(log.tokens_used.get("total_tokens", 0) for log in logs) + + return { + "total_cost": total_cost, + "cost_by_agent": dict(agent_costs), + "cost_by_task_type": dict(task_type_costs), + "cost_per_token": total_cost / total_tokens if total_tokens > 0 else 0, + "top_cost_drivers": sorted(task_type_costs.items(), key=lambda x: x[1], reverse=True)[:5] + } + + def _check_sla_compliance(self, metrics: PerformanceMetrics) -> Dict[str, Any]: + """Check SLA compliance""" + thresholds = self.performance_thresholds + + compliance = { + "success_rate": { + "target": 0.95, + "actual": metrics.success_rate, + "compliant": metrics.success_rate >= 0.95, + "gap": max(0, 0.95 - metrics.success_rate) + }, + "average_latency": { + "target": 10000, # 10 seconds + "actual": metrics.average_duration_ms, + "compliant": metrics.average_duration_ms <= 10000, + "gap": max(0, metrics.average_duration_ms - 10000) + }, + "error_rate": { + "target": 0.05, # 5% + "actual": metrics.error_rate, + "compliant": metrics.error_rate <= 0.05, + "gap": max(0, metrics.error_rate - 0.05) + } + } + + overall_compliance = all(sla["compliant"] for sla in compliance.values()) + + return { + "overall_compliant": overall_compliance, + "sla_details": compliance, + "compliance_score": sum(1 for sla in compliance.values() if sla["compliant"]) / len(compliance) + } + + def _assess_overall_health(self, metrics: PerformanceMetrics) -> str: + """Assess overall system health""" + health_score = 0 + + # Success rate contribution (40%) + if metrics.success_rate >= 0.95: + health_score += 40 + elif metrics.success_rate >= 0.90: + health_score += 30 + elif metrics.success_rate >= 0.80: + health_score += 20 + else: + health_score += 10 + + # Performance contribution (30%) + if metrics.average_duration_ms <= 5000: + health_score += 30 + elif metrics.average_duration_ms <= 10000: + health_score += 20 + elif metrics.average_duration_ms <= 30000: + health_score += 15 + else: + health_score += 5 + + # Error rate contribution (20%) + if metrics.error_rate <= 0.02: + health_score += 20 + elif metrics.error_rate <= 0.05: + health_score += 15 + elif metrics.error_rate <= 0.10: + health_score += 10 + else: + health_score += 0 + + # Cost efficiency contribution (10%) + if metrics.cost_per_token <= 0.00005: + health_score += 10 + elif metrics.cost_per_token <= 0.0001: + health_score += 7 + else: + health_score += 3 + + if health_score >= 85: + return "excellent" + elif health_score >= 70: + return "good" + elif health_score >= 50: + return "fair" + else: + return "poor" + + def _extract_key_findings(self, metrics: PerformanceMetrics, + errors: List[ErrorAnalysis], + bottlenecks: List[BottleneckAnalysis]) -> List[str]: + """Extract key findings from analysis""" + findings = [] + + # Performance findings + if metrics.success_rate < 0.9: + findings.append(f"Success rate ({metrics.success_rate:.1%}) below target") + + if metrics.average_duration_ms > 15000: + findings.append(f"High average latency ({metrics.average_duration_ms/1000:.1f}s)") + + # Error findings + high_impact_errors = [e for e in errors if e.impact_level == "high"] + if high_impact_errors: + findings.append(f"{len(high_impact_errors)} high-impact error patterns identified") + + # Bottleneck findings + critical_bottlenecks = [b for b in bottlenecks if b.severity == "critical"] + if critical_bottlenecks: + findings.append(f"{len(critical_bottlenecks)} critical bottlenecks found") + + # Cost findings + if metrics.cost_per_token > 0.0001: + findings.append("Token usage costs above optimal range") + + return findings + + +def main(): + parser = argparse.ArgumentParser(description="Multi-Agent System Performance Evaluator") + parser.add_argument("input_file", help="JSON file with execution logs") + parser.add_argument("-o", "--output", help="Output file prefix (default: evaluation_report)") + parser.add_argument("--format", choices=["json", "both"], default="both", + help="Output format") + parser.add_argument("--detailed", action="store_true", + help="Include detailed analysis in output") + + args = parser.parse_args() + + try: + # Load execution logs + with open(args.input_file, 'r') as f: + logs_data = json.load(f) + + # Parse logs + evaluator = AgentEvaluator() + logs = evaluator.parse_execution_logs(logs_data.get("execution_logs", [])) + + if not logs: + print("No valid execution logs found in input file", file=sys.stderr) + sys.exit(1) + + # Generate evaluation report + report = evaluator.generate_report(logs) + + # Prepare output + output_data = asdict(report) + + # Output files + output_prefix = args.output or "evaluation_report" + + if args.format in ["json", "both"]: + with open(f"{output_prefix}.json", 'w') as f: + json.dump(output_data, f, indent=2, default=str) + print(f"JSON report written to {output_prefix}.json") + + if args.format == "both": + # Generate separate detailed files + + # Performance summary + summary_data = { + "summary": report.summary, + "system_metrics": asdict(report.system_metrics), + "sla_compliance": report.sla_compliance + } + with open(f"{output_prefix}_summary.json", 'w') as f: + json.dump(summary_data, f, indent=2, default=str) + print(f"Summary report written to {output_prefix}_summary.json") + + # Recommendations + recommendations_data = { + "optimization_recommendations": [asdict(rec) for rec in report.optimization_recommendations], + "bottleneck_analysis": [asdict(b) for b in report.bottleneck_analysis] + } + with open(f"{output_prefix}_recommendations.json", 'w') as f: + json.dump(recommendations_data, f, indent=2) + print(f"Recommendations written to {output_prefix}_recommendations.json") + + # Error analysis + error_data = { + "error_analysis": [asdict(e) for e in report.error_analysis], + "error_summary": { + "total_errors": sum(e.count for e in report.error_analysis), + "high_impact_errors": len([e for e in report.error_analysis if e.impact_level == "high"]) + } + } + with open(f"{output_prefix}_errors.json", 'w') as f: + json.dump(error_data, f, indent=2) + print(f"Error analysis written to {output_prefix}_errors.json") + + # Print executive summary + print(f"\n{'='*60}") + print(f"AGENT SYSTEM EVALUATION REPORT") + print(f"{'='*60}") + print(f"Overall Health: {report.summary['overall_health'].upper()}") + print(f"Total Tasks: {report.system_metrics.total_tasks}") + print(f"Success Rate: {report.system_metrics.success_rate:.1%}") + print(f"Average Duration: {report.system_metrics.average_duration_ms/1000:.1f}s") + print(f"Total Cost: ${report.system_metrics.total_cost_usd:.2f}") + print(f"Agents Analyzed: {len(report.agent_metrics)}") + + print(f"\nKey Findings:") + for finding in report.summary['key_findings']: + print(f" • {finding}") + + print(f"\nTop Recommendations:") + high_priority_recs = [r for r in report.optimization_recommendations if r.priority == "high"][:3] + for i, rec in enumerate(high_priority_recs, 1): + print(f" {i}. {rec.title}") + + if report.summary['critical_issues'] > 0: + print(f"\n⚠️ CRITICAL: {report.summary['critical_issues']} critical issues require immediate attention") + + print(f"\n📊 Detailed reports available in generated files") + print(f"{'='*60}") + + except Exception as e: + print(f"Error: {e}", file=sys.stderr) + sys.exit(1) + + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/skills/agent-designer/agent_planner.py b/skills/agent-designer/agent_planner.py new file mode 100644 index 00000000..46b8aed6 --- /dev/null +++ b/skills/agent-designer/agent_planner.py @@ -0,0 +1,911 @@ +#!/usr/bin/env python3 +""" +Agent Planner - Multi-Agent System Architecture Designer + +Given a system description (goal, tasks, constraints, team size), designs a multi-agent +architecture: defines agent roles, responsibilities, capabilities needed, communication +topology, tool requirements. Generates architecture diagram (Mermaid). + +Input: system requirements JSON +Output: agent architecture + role definitions + Mermaid diagram + implementation roadmap +""" + +import json +import argparse +import sys +from typing import Dict, List, Any, Optional, Tuple +from dataclasses import dataclass, asdict +from enum import Enum + + +class AgentArchitecturePattern(Enum): + """Supported agent architecture patterns""" + SINGLE_AGENT = "single_agent" + SUPERVISOR = "supervisor" + SWARM = "swarm" + HIERARCHICAL = "hierarchical" + PIPELINE = "pipeline" + + +class CommunicationPattern(Enum): + """Agent communication patterns""" + DIRECT_MESSAGE = "direct_message" + SHARED_STATE = "shared_state" + EVENT_DRIVEN = "event_driven" + MESSAGE_QUEUE = "message_queue" + + +class AgentRole(Enum): + """Standard agent role archetypes""" + COORDINATOR = "coordinator" + SPECIALIST = "specialist" + INTERFACE = "interface" + MONITOR = "monitor" + + +@dataclass +class Tool: + """Tool definition for agents""" + name: str + description: str + input_schema: Dict[str, Any] + output_schema: Dict[str, Any] + capabilities: List[str] + reliability: str = "high" # high, medium, low + latency: str = "low" # low, medium, high + + +@dataclass +class AgentDefinition: + """Complete agent definition""" + name: str + role: str + archetype: AgentRole + responsibilities: List[str] + capabilities: List[str] + tools: List[Tool] + communication_interfaces: List[str] + constraints: Dict[str, Any] + success_criteria: List[str] + dependencies: List[str] = None + + +@dataclass +class CommunicationLink: + """Communication link between agents""" + from_agent: str + to_agent: str + pattern: CommunicationPattern + data_format: str + frequency: str + criticality: str + + +@dataclass +class SystemRequirements: + """Input system requirements""" + goal: str + description: str + tasks: List[str] + constraints: Dict[str, Any] + team_size: int + performance_requirements: Dict[str, Any] + safety_requirements: List[str] + integration_requirements: List[str] + scale_requirements: Dict[str, Any] + + +@dataclass +class ArchitectureDesign: + """Complete architecture design output""" + pattern: AgentArchitecturePattern + agents: List[AgentDefinition] + communication_topology: List[CommunicationLink] + shared_resources: List[Dict[str, Any]] + guardrails: List[Dict[str, Any]] + scaling_strategy: Dict[str, Any] + failure_handling: Dict[str, Any] + + +class AgentPlanner: + """Multi-agent system architecture planner""" + + def __init__(self): + self.common_tools = self._define_common_tools() + self.pattern_heuristics = self._define_pattern_heuristics() + + def _define_common_tools(self) -> Dict[str, Tool]: + """Define commonly used tools across agents""" + return { + "web_search": Tool( + name="web_search", + description="Search the web for information", + input_schema={"type": "object", "properties": {"query": {"type": "string"}}}, + output_schema={"type": "object", "properties": {"results": {"type": "array"}}}, + capabilities=["research", "information_gathering"], + reliability="high", + latency="medium" + ), + "code_executor": Tool( + name="code_executor", + description="Execute code in various languages", + input_schema={"type": "object", "properties": {"language": {"type": "string"}, "code": {"type": "string"}}}, + output_schema={"type": "object", "properties": {"result": {"type": "string"}, "error": {"type": "string"}}}, + capabilities=["code_execution", "testing", "automation"], + reliability="high", + latency="low" + ), + "file_manager": Tool( + name="file_manager", + description="Manage files and directories", + input_schema={"type": "object", "properties": {"action": {"type": "string"}, "path": {"type": "string"}}}, + output_schema={"type": "object", "properties": {"success": {"type": "boolean"}, "content": {"type": "string"}}}, + capabilities=["file_operations", "data_management"], + reliability="high", + latency="low" + ), + "data_analyzer": Tool( + name="data_analyzer", + description="Analyze and process data", + input_schema={"type": "object", "properties": {"data": {"type": "object"}, "analysis_type": {"type": "string"}}}, + output_schema={"type": "object", "properties": {"insights": {"type": "array"}, "metrics": {"type": "object"}}}, + capabilities=["data_analysis", "statistics", "visualization"], + reliability="high", + latency="medium" + ), + "api_client": Tool( + name="api_client", + description="Make API calls to external services", + input_schema={"type": "object", "properties": {"url": {"type": "string"}, "method": {"type": "string"}, "data": {"type": "object"}}}, + output_schema={"type": "object", "properties": {"response": {"type": "object"}, "status": {"type": "integer"}}}, + capabilities=["integration", "external_services"], + reliability="medium", + latency="medium" + ) + } + + def _define_pattern_heuristics(self) -> Dict[AgentArchitecturePattern, Dict[str, Any]]: + """Define heuristics for selecting architecture patterns""" + return { + AgentArchitecturePattern.SINGLE_AGENT: { + "team_size_range": (1, 1), + "task_complexity": "simple", + "coordination_overhead": "none", + "suitable_for": ["simple tasks", "prototyping", "single domain"], + "scaling_limit": "low" + }, + AgentArchitecturePattern.SUPERVISOR: { + "team_size_range": (2, 8), + "task_complexity": "medium", + "coordination_overhead": "low", + "suitable_for": ["hierarchical tasks", "clear delegation", "quality control"], + "scaling_limit": "medium" + }, + AgentArchitecturePattern.SWARM: { + "team_size_range": (3, 20), + "task_complexity": "high", + "coordination_overhead": "high", + "suitable_for": ["parallel processing", "distributed problem solving", "fault tolerance"], + "scaling_limit": "high" + }, + AgentArchitecturePattern.HIERARCHICAL: { + "team_size_range": (5, 50), + "task_complexity": "very high", + "coordination_overhead": "medium", + "suitable_for": ["large organizations", "complex workflows", "enterprise systems"], + "scaling_limit": "very high" + }, + AgentArchitecturePattern.PIPELINE: { + "team_size_range": (3, 15), + "task_complexity": "medium", + "coordination_overhead": "low", + "suitable_for": ["sequential processing", "data pipelines", "assembly line tasks"], + "scaling_limit": "medium" + } + } + + def select_architecture_pattern(self, requirements: SystemRequirements) -> AgentArchitecturePattern: + """Select the most appropriate architecture pattern based on requirements""" + team_size = requirements.team_size + task_count = len(requirements.tasks) + performance_reqs = requirements.performance_requirements + + # Score each pattern based on requirements + pattern_scores = {} + + for pattern, heuristics in self.pattern_heuristics.items(): + score = 0 + + # Team size fit + min_size, max_size = heuristics["team_size_range"] + if min_size <= team_size <= max_size: + score += 3 + elif abs(team_size - min_size) <= 2 or abs(team_size - max_size) <= 2: + score += 1 + + # Task complexity assessment + complexity_indicators = [ + "parallel" in requirements.description.lower(), + "sequential" in requirements.description.lower(), + "hierarchical" in requirements.description.lower(), + "distributed" in requirements.description.lower(), + task_count > 5, + len(requirements.constraints) > 3 + ] + + complexity_score = sum(complexity_indicators) + + if pattern == AgentArchitecturePattern.SINGLE_AGENT and complexity_score <= 2: + score += 2 + elif pattern == AgentArchitecturePattern.SUPERVISOR and 2 <= complexity_score <= 4: + score += 2 + elif pattern == AgentArchitecturePattern.PIPELINE and "sequential" in requirements.description.lower(): + score += 3 + elif pattern == AgentArchitecturePattern.SWARM and "parallel" in requirements.description.lower(): + score += 3 + elif pattern == AgentArchitecturePattern.HIERARCHICAL and complexity_score >= 4: + score += 2 + + # Performance requirements + if performance_reqs.get("high_throughput", False) and pattern in [AgentArchitecturePattern.SWARM, AgentArchitecturePattern.PIPELINE]: + score += 2 + if performance_reqs.get("fault_tolerance", False) and pattern == AgentArchitecturePattern.SWARM: + score += 2 + if performance_reqs.get("low_latency", False) and pattern in [AgentArchitecturePattern.SINGLE_AGENT, AgentArchitecturePattern.PIPELINE]: + score += 1 + + pattern_scores[pattern] = score + + # Select the highest scoring pattern + best_pattern = max(pattern_scores.items(), key=lambda x: x[1])[0] + return best_pattern + + def design_agents(self, requirements: SystemRequirements, pattern: AgentArchitecturePattern) -> List[AgentDefinition]: + """Design individual agents based on requirements and architecture pattern""" + agents = [] + + if pattern == AgentArchitecturePattern.SINGLE_AGENT: + agents = self._design_single_agent(requirements) + elif pattern == AgentArchitecturePattern.SUPERVISOR: + agents = self._design_supervisor_agents(requirements) + elif pattern == AgentArchitecturePattern.SWARM: + agents = self._design_swarm_agents(requirements) + elif pattern == AgentArchitecturePattern.HIERARCHICAL: + agents = self._design_hierarchical_agents(requirements) + elif pattern == AgentArchitecturePattern.PIPELINE: + agents = self._design_pipeline_agents(requirements) + + return agents + + def _design_single_agent(self, requirements: SystemRequirements) -> List[AgentDefinition]: + """Design a single general-purpose agent""" + all_tools = list(self.common_tools.values()) + + agent = AgentDefinition( + name="universal_agent", + role="Universal Task Handler", + archetype=AgentRole.SPECIALIST, + responsibilities=requirements.tasks, + capabilities=["general_purpose", "multi_domain", "adaptable"], + tools=all_tools, + communication_interfaces=["direct_user_interface"], + constraints={ + "max_concurrent_tasks": 1, + "memory_limit": "high", + "response_time": "fast" + }, + success_criteria=["complete all assigned tasks", "maintain quality standards", "respond within time limits"], + dependencies=[] + ) + + return [agent] + + def _design_supervisor_agents(self, requirements: SystemRequirements) -> List[AgentDefinition]: + """Design supervisor pattern agents""" + agents = [] + + # Create supervisor agent + supervisor = AgentDefinition( + name="supervisor_agent", + role="Task Coordinator and Quality Controller", + archetype=AgentRole.COORDINATOR, + responsibilities=[ + "task_decomposition", + "delegation", + "progress_monitoring", + "quality_assurance", + "result_aggregation" + ], + capabilities=["planning", "coordination", "evaluation", "decision_making"], + tools=[self.common_tools["file_manager"], self.common_tools["data_analyzer"]], + communication_interfaces=["user_interface", "agent_messaging"], + constraints={ + "max_concurrent_supervisions": 5, + "decision_timeout": "30s" + }, + success_criteria=["successful task completion", "optimal resource utilization", "quality standards met"], + dependencies=[] + ) + agents.append(supervisor) + + # Create specialist agents based on task domains + task_domains = self._identify_task_domains(requirements.tasks) + for i, domain in enumerate(task_domains[:requirements.team_size - 1]): + specialist = AgentDefinition( + name=f"{domain}_specialist", + role=f"{domain.title()} Specialist", + archetype=AgentRole.SPECIALIST, + responsibilities=[task for task in requirements.tasks if domain in task.lower()], + capabilities=[f"{domain}_expertise", "specialized_tools", "domain_knowledge"], + tools=self._select_tools_for_domain(domain), + communication_interfaces=["supervisor_messaging"], + constraints={ + "domain_scope": domain, + "task_queue_size": 10 + }, + success_criteria=[f"excel in {domain} tasks", "maintain domain expertise", "provide quality output"], + dependencies=["supervisor_agent"] + ) + agents.append(specialist) + + return agents + + def _design_swarm_agents(self, requirements: SystemRequirements) -> List[AgentDefinition]: + """Design swarm pattern agents""" + agents = [] + + # Create peer agents with overlapping capabilities + agent_count = min(requirements.team_size, 10) # Reasonable swarm size + base_capabilities = ["collaboration", "consensus", "adaptation", "peer_communication"] + + for i in range(agent_count): + agent = AgentDefinition( + name=f"swarm_agent_{i+1}", + role=f"Collaborative Worker #{i+1}", + archetype=AgentRole.SPECIALIST, + responsibilities=requirements.tasks, # All agents can handle all tasks + capabilities=base_capabilities + [f"specialization_{i%3}"], # Some specialization + tools=list(self.common_tools.values()), + communication_interfaces=["peer_messaging", "broadcast", "consensus_protocol"], + constraints={ + "peer_discovery_timeout": "10s", + "consensus_threshold": 0.6, + "max_retries": 3 + }, + success_criteria=["contribute to group goals", "maintain peer relationships", "adapt to failures"], + dependencies=[f"swarm_agent_{j+1}" for j in range(agent_count) if j != i] + ) + agents.append(agent) + + return agents + + def _design_hierarchical_agents(self, requirements: SystemRequirements) -> List[AgentDefinition]: + """Design hierarchical pattern agents""" + agents = [] + + # Create management hierarchy + levels = min(3, requirements.team_size // 3) # Reasonable hierarchy depth + agents_per_level = requirements.team_size // levels + + # Top level manager + manager = AgentDefinition( + name="executive_manager", + role="Executive Manager", + archetype=AgentRole.COORDINATOR, + responsibilities=["strategic_planning", "resource_allocation", "performance_monitoring"], + capabilities=["leadership", "strategy", "resource_management", "oversight"], + tools=[self.common_tools["data_analyzer"], self.common_tools["file_manager"]], + communication_interfaces=["executive_dashboard", "management_messaging"], + constraints={"management_span": 5, "decision_authority": "high"}, + success_criteria=["achieve system goals", "optimize resource usage", "maintain quality"], + dependencies=[] + ) + agents.append(manager) + + # Middle managers + for i in range(agents_per_level - 1): + middle_manager = AgentDefinition( + name=f"team_manager_{i+1}", + role=f"Team Manager #{i+1}", + archetype=AgentRole.COORDINATOR, + responsibilities=["team_coordination", "task_distribution", "progress_tracking"], + capabilities=["team_management", "coordination", "reporting"], + tools=[self.common_tools["file_manager"]], + communication_interfaces=["management_messaging", "team_messaging"], + constraints={"team_size": 3, "reporting_frequency": "hourly"}, + success_criteria=["team performance", "task completion", "team satisfaction"], + dependencies=["executive_manager"] + ) + agents.append(middle_manager) + + # Workers + remaining_agents = requirements.team_size - len(agents) + for i in range(remaining_agents): + worker = AgentDefinition( + name=f"worker_agent_{i+1}", + role=f"Task Worker #{i+1}", + archetype=AgentRole.SPECIALIST, + responsibilities=["task_execution", "result_delivery", "status_reporting"], + capabilities=["task_execution", "specialized_skills", "reliability"], + tools=self._select_diverse_tools(), + communication_interfaces=["team_messaging"], + constraints={"task_focus": "single", "reporting_interval": "30min"}, + success_criteria=["complete assigned tasks", "maintain quality", "meet deadlines"], + dependencies=[f"team_manager_{(i // 3) + 1}"] + ) + agents.append(worker) + + return agents + + def _design_pipeline_agents(self, requirements: SystemRequirements) -> List[AgentDefinition]: + """Design pipeline pattern agents""" + agents = [] + + # Create sequential processing stages + pipeline_stages = self._identify_pipeline_stages(requirements.tasks) + + for i, stage in enumerate(pipeline_stages): + agent = AgentDefinition( + name=f"pipeline_stage_{i+1}_{stage}", + role=f"Pipeline Stage {i+1}: {stage.title()}", + archetype=AgentRole.SPECIALIST, + responsibilities=[f"process_{stage}", f"validate_{stage}_output", "handoff_to_next_stage"], + capabilities=[f"{stage}_processing", "quality_control", "data_transformation"], + tools=self._select_tools_for_stage(stage), + communication_interfaces=["pipeline_queue", "stage_messaging"], + constraints={ + "processing_order": i + 1, + "batch_size": 10, + "stage_timeout": "5min" + }, + success_criteria=[f"successfully process {stage}", "maintain data integrity", "meet throughput targets"], + dependencies=[f"pipeline_stage_{i}_{pipeline_stages[i-1]}"] if i > 0 else [] + ) + agents.append(agent) + + return agents + + def _identify_task_domains(self, tasks: List[str]) -> List[str]: + """Identify distinct domains from task list""" + domains = [] + domain_keywords = { + "research": ["research", "search", "find", "investigate", "analyze"], + "development": ["code", "build", "develop", "implement", "program"], + "data": ["data", "process", "analyze", "calculate", "compute"], + "communication": ["write", "send", "message", "communicate", "report"], + "file": ["file", "document", "save", "load", "manage"] + } + + for domain, keywords in domain_keywords.items(): + if any(keyword in " ".join(tasks).lower() for keyword in keywords): + domains.append(domain) + + return domains[:5] # Limit to 5 domains + + def _identify_pipeline_stages(self, tasks: List[str]) -> List[str]: + """Identify pipeline stages from task list""" + # Common pipeline patterns + common_stages = ["input", "process", "transform", "validate", "output"] + + # Try to infer stages from tasks + stages = [] + task_text = " ".join(tasks).lower() + + if "collect" in task_text or "gather" in task_text: + stages.append("collection") + if "process" in task_text or "transform" in task_text: + stages.append("processing") + if "analyze" in task_text or "evaluate" in task_text: + stages.append("analysis") + if "validate" in task_text or "check" in task_text: + stages.append("validation") + if "output" in task_text or "deliver" in task_text or "report" in task_text: + stages.append("output") + + # Default to common stages if none identified + return stages if stages else common_stages[:min(5, len(tasks))] + + def _select_tools_for_domain(self, domain: str) -> List[Tool]: + """Select appropriate tools for a specific domain""" + domain_tools = { + "research": [self.common_tools["web_search"], self.common_tools["data_analyzer"]], + "development": [self.common_tools["code_executor"], self.common_tools["file_manager"]], + "data": [self.common_tools["data_analyzer"], self.common_tools["file_manager"]], + "communication": [self.common_tools["api_client"], self.common_tools["file_manager"]], + "file": [self.common_tools["file_manager"]] + } + + return domain_tools.get(domain, [self.common_tools["api_client"]]) + + def _select_tools_for_stage(self, stage: str) -> List[Tool]: + """Select appropriate tools for a pipeline stage""" + stage_tools = { + "input": [self.common_tools["api_client"], self.common_tools["file_manager"]], + "collection": [self.common_tools["web_search"], self.common_tools["api_client"]], + "process": [self.common_tools["code_executor"], self.common_tools["data_analyzer"]], + "processing": [self.common_tools["data_analyzer"], self.common_tools["code_executor"]], + "transform": [self.common_tools["data_analyzer"], self.common_tools["code_executor"]], + "analysis": [self.common_tools["data_analyzer"]], + "validate": [self.common_tools["data_analyzer"]], + "validation": [self.common_tools["data_analyzer"]], + "output": [self.common_tools["file_manager"], self.common_tools["api_client"]] + } + + return stage_tools.get(stage, [self.common_tools["file_manager"]]) + + def _select_diverse_tools(self) -> List[Tool]: + """Select a diverse set of tools for general purpose agents""" + return [ + self.common_tools["file_manager"], + self.common_tools["code_executor"], + self.common_tools["data_analyzer"] + ] + + def design_communication_topology(self, agents: List[AgentDefinition], pattern: AgentArchitecturePattern) -> List[CommunicationLink]: + """Design communication links between agents""" + links = [] + + if pattern == AgentArchitecturePattern.SINGLE_AGENT: + # No inter-agent communication needed + return [] + + elif pattern == AgentArchitecturePattern.SUPERVISOR: + supervisor = next(agent for agent in agents if agent.archetype == AgentRole.COORDINATOR) + specialists = [agent for agent in agents if agent.archetype == AgentRole.SPECIALIST] + + for specialist in specialists: + # Bidirectional communication with supervisor + links.append(CommunicationLink( + from_agent=supervisor.name, + to_agent=specialist.name, + pattern=CommunicationPattern.DIRECT_MESSAGE, + data_format="json", + frequency="on_demand", + criticality="high" + )) + links.append(CommunicationLink( + from_agent=specialist.name, + to_agent=supervisor.name, + pattern=CommunicationPattern.DIRECT_MESSAGE, + data_format="json", + frequency="on_completion", + criticality="high" + )) + + elif pattern == AgentArchitecturePattern.SWARM: + # All-to-all communication for swarm + for i, agent1 in enumerate(agents): + for j, agent2 in enumerate(agents): + if i != j: + links.append(CommunicationLink( + from_agent=agent1.name, + to_agent=agent2.name, + pattern=CommunicationPattern.EVENT_DRIVEN, + data_format="json", + frequency="periodic", + criticality="medium" + )) + + elif pattern == AgentArchitecturePattern.HIERARCHICAL: + # Hierarchical communication based on dependencies + for agent in agents: + if agent.dependencies: + for dependency in agent.dependencies: + links.append(CommunicationLink( + from_agent=dependency, + to_agent=agent.name, + pattern=CommunicationPattern.DIRECT_MESSAGE, + data_format="json", + frequency="scheduled", + criticality="high" + )) + links.append(CommunicationLink( + from_agent=agent.name, + to_agent=dependency, + pattern=CommunicationPattern.DIRECT_MESSAGE, + data_format="json", + frequency="on_completion", + criticality="high" + )) + + elif pattern == AgentArchitecturePattern.PIPELINE: + # Sequential pipeline communication + for i in range(len(agents) - 1): + links.append(CommunicationLink( + from_agent=agents[i].name, + to_agent=agents[i + 1].name, + pattern=CommunicationPattern.MESSAGE_QUEUE, + data_format="json", + frequency="continuous", + criticality="high" + )) + + return links + + def generate_mermaid_diagram(self, design: ArchitectureDesign) -> str: + """Generate Mermaid diagram for the architecture""" + diagram = ["graph TD"] + + # Add agent nodes + for agent in design.agents: + node_style = self._get_node_style(agent.archetype) + diagram.append(f" {agent.name}[{agent.role}]{node_style}") + + # Add communication links + for link in design.communication_topology: + arrow_style = self._get_arrow_style(link.pattern, link.criticality) + diagram.append(f" {link.from_agent} {arrow_style} {link.to_agent}") + + # Add styling + diagram.extend([ + "", + " classDef coordinator fill:#e1f5fe,stroke:#01579b,stroke-width:2px", + " classDef specialist fill:#f3e5f5,stroke:#4a148c,stroke-width:2px", + " classDef interface fill:#e8f5e8,stroke:#1b5e20,stroke-width:2px", + " classDef monitor fill:#fff3e0,stroke:#e65100,stroke-width:2px" + ]) + + # Apply classes to nodes + for agent in design.agents: + class_name = agent.archetype.value + diagram.append(f" class {agent.name} {class_name}") + + return "\n".join(diagram) + + def _get_node_style(self, archetype: AgentRole) -> str: + """Get node styling based on archetype""" + styles = { + AgentRole.COORDINATOR: ":::coordinator", + AgentRole.SPECIALIST: ":::specialist", + AgentRole.INTERFACE: ":::interface", + AgentRole.MONITOR: ":::monitor" + } + return styles.get(archetype, "") + + def _get_arrow_style(self, pattern: CommunicationPattern, criticality: str) -> str: + """Get arrow styling based on communication pattern and criticality""" + base_arrows = { + CommunicationPattern.DIRECT_MESSAGE: "-->", + CommunicationPattern.SHARED_STATE: "-.->", + CommunicationPattern.EVENT_DRIVEN: "===>", + CommunicationPattern.MESSAGE_QUEUE: "===" + } + + arrow = base_arrows.get(pattern, "-->") + + # Modify for criticality + if criticality == "high": + return arrow + elif criticality == "medium": + return arrow.replace("-", ".") + else: + return arrow.replace("-", ":") + + def generate_implementation_roadmap(self, design: ArchitectureDesign, requirements: SystemRequirements) -> Dict[str, Any]: + """Generate implementation roadmap""" + phases = [] + + # Phase 1: Core Infrastructure + phases.append({ + "phase": 1, + "name": "Core Infrastructure", + "duration": "2-3 weeks", + "tasks": [ + "Set up development environment", + "Implement basic agent framework", + "Create communication infrastructure", + "Set up monitoring and logging", + "Implement basic tools" + ], + "deliverables": [ + "Agent runtime framework", + "Communication layer", + "Basic monitoring dashboard" + ] + }) + + # Phase 2: Agent Implementation + phases.append({ + "phase": 2, + "name": "Agent Implementation", + "duration": "3-4 weeks", + "tasks": [ + "Implement individual agent logic", + "Create agent-specific tools", + "Implement communication protocols", + "Add error handling and recovery", + "Create agent configuration system" + ], + "deliverables": [ + "Functional agent implementations", + "Tool integration", + "Configuration management" + ] + }) + + # Phase 3: Integration and Testing + phases.append({ + "phase": 3, + "name": "Integration and Testing", + "duration": "2-3 weeks", + "tasks": [ + "Integrate all agents", + "End-to-end testing", + "Performance optimization", + "Security implementation", + "Documentation creation" + ], + "deliverables": [ + "Integrated system", + "Test suite", + "Performance benchmarks", + "Security audit report" + ] + }) + + # Phase 4: Deployment and Monitoring + phases.append({ + "phase": 4, + "name": "Deployment and Monitoring", + "duration": "1-2 weeks", + "tasks": [ + "Production deployment", + "Monitoring setup", + "Alerting configuration", + "User training", + "Go-live support" + ], + "deliverables": [ + "Production system", + "Monitoring dashboard", + "Operational runbooks", + "Training materials" + ] + }) + + return { + "total_duration": "8-12 weeks", + "phases": phases, + "critical_path": [ + "Agent framework implementation", + "Communication layer development", + "Integration testing", + "Production deployment" + ], + "risks": [ + { + "risk": "Communication complexity", + "impact": "high", + "mitigation": "Start with simple protocols, iterate" + }, + { + "risk": "Agent coordination failures", + "impact": "medium", + "mitigation": "Implement robust error handling and fallbacks" + }, + { + "risk": "Performance bottlenecks", + "impact": "medium", + "mitigation": "Early performance testing and optimization" + } + ], + "success_criteria": requirements.safety_requirements + [ + "All agents operational", + "Communication working reliably", + "Performance targets met", + "Error rate below 1%" + ] + } + + def plan_system(self, requirements: SystemRequirements) -> Tuple[ArchitectureDesign, str, Dict[str, Any]]: + """Main planning function""" + # Select architecture pattern + pattern = self.select_architecture_pattern(requirements) + + # Design agents + agents = self.design_agents(requirements, pattern) + + # Design communication topology + communication_topology = self.design_communication_topology(agents, pattern) + + # Create complete design + design = ArchitectureDesign( + pattern=pattern, + agents=agents, + communication_topology=communication_topology, + shared_resources=[ + {"type": "message_queue", "capacity": 1000}, + {"type": "shared_memory", "size": "1GB"}, + {"type": "event_store", "retention": "30 days"} + ], + guardrails=[ + {"type": "input_validation", "rules": "strict_schema_enforcement"}, + {"type": "rate_limiting", "limit": "100_requests_per_minute"}, + {"type": "output_filtering", "rules": "content_safety_check"} + ], + scaling_strategy={ + "horizontal_scaling": True, + "auto_scaling_triggers": ["cpu > 80%", "queue_depth > 100"], + "max_instances_per_agent": 5 + }, + failure_handling={ + "retry_policy": "exponential_backoff", + "circuit_breaker": True, + "fallback_strategies": ["graceful_degradation", "human_escalation"] + } + ) + + # Generate Mermaid diagram + mermaid_diagram = self.generate_mermaid_diagram(design) + + # Generate implementation roadmap + roadmap = self.generate_implementation_roadmap(design, requirements) + + return design, mermaid_diagram, roadmap + + +def main(): + parser = argparse.ArgumentParser(description="Multi-Agent System Architecture Planner") + parser.add_argument("input_file", help="JSON file with system requirements") + parser.add_argument("-o", "--output", help="Output file prefix (default: agent_architecture)") + parser.add_argument("--format", choices=["json", "yaml", "both"], default="both", + help="Output format") + + args = parser.parse_args() + + try: + # Load requirements + with open(args.input_file, 'r') as f: + requirements_data = json.load(f) + + requirements = SystemRequirements(**requirements_data) + + # Plan the system + planner = AgentPlanner() + design, mermaid_diagram, roadmap = planner.plan_system(requirements) + + # Prepare output + output_data = { + "architecture_design": asdict(design), + "mermaid_diagram": mermaid_diagram, + "implementation_roadmap": roadmap, + "metadata": { + "generated_by": "agent_planner.py", + "requirements_file": args.input_file, + "architecture_pattern": design.pattern.value, + "agent_count": len(design.agents) + } + } + + # Output files + output_prefix = args.output or "agent_architecture" + + if args.format in ["json", "both"]: + with open(f"{output_prefix}.json", 'w') as f: + json.dump(output_data, f, indent=2, default=str) + print(f"JSON output written to {output_prefix}.json") + + if args.format in ["both"]: + # Also create separate files for key components + with open(f"{output_prefix}_diagram.mmd", 'w') as f: + f.write(mermaid_diagram) + print(f"Mermaid diagram written to {output_prefix}_diagram.mmd") + + with open(f"{output_prefix}_roadmap.json", 'w') as f: + json.dump(roadmap, f, indent=2) + print(f"Implementation roadmap written to {output_prefix}_roadmap.json") + + # Print summary + print(f"\nArchitecture Summary:") + print(f"Pattern: {design.pattern.value}") + print(f"Agents: {len(design.agents)}") + print(f"Communication Links: {len(design.communication_topology)}") + print(f"Estimated Duration: {roadmap['total_duration']}") + + except Exception as e: + print(f"Error: {e}", file=sys.stderr) + sys.exit(1) + + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/skills/agent-designer/assets/sample_execution_logs.json b/skills/agent-designer/assets/sample_execution_logs.json new file mode 100644 index 00000000..13ec29bd --- /dev/null +++ b/skills/agent-designer/assets/sample_execution_logs.json @@ -0,0 +1,543 @@ +{ + "execution_logs": [ + { + "task_id": "task_001", + "agent_id": "research_agent_1", + "task_type": "web_research", + "task_description": "Research recent developments in artificial intelligence", + "start_time": "2024-01-15T09:00:00Z", + "end_time": "2024-01-15T09:02:34Z", + "duration_ms": 154000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2300, + "success": true, + "parameters": { + "query": "artificial intelligence developments 2024", + "limit": 10 + } + }, + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2100, + "success": true, + "parameters": { + "query": "machine learning breakthroughs recent", + "limit": 5 + } + }, + { + "type": "analysis", + "description": "Synthesize search results", + "duration_ms": 149600, + "success": true + } + ], + "results": { + "summary": "Found 15 relevant sources covering recent AI developments including GPT-4 improvements, autonomous vehicle progress, and medical AI applications.", + "sources_found": 15, + "quality_score": 0.92 + }, + "tokens_used": { + "input_tokens": 1250, + "output_tokens": 2800, + "total_tokens": 4050 + }, + "cost_usd": 0.081, + "error_details": null, + "tools_used": ["web_search"], + "retry_count": 0, + "metadata": { + "user_id": "user_123", + "session_id": "session_abc", + "request_priority": "normal" + } + }, + { + "task_id": "task_002", + "agent_id": "data_agent_1", + "task_type": "data_analysis", + "task_description": "Analyze sales performance data for Q4 2023", + "start_time": "2024-01-15T09:05:00Z", + "end_time": "2024-01-15T09:07:45Z", + "duration_ms": 165000, + "status": "success", + "actions": [ + { + "type": "data_ingestion", + "description": "Load Q4 sales data", + "duration_ms": 5000, + "success": true + }, + { + "type": "tool_call", + "tool_name": "data_analyzer", + "duration_ms": 155000, + "success": true, + "parameters": { + "analysis_type": "descriptive", + "target_column": "revenue" + } + }, + { + "type": "visualization", + "description": "Generate charts and graphs", + "duration_ms": 5000, + "success": true + } + ], + "results": { + "insights": [ + "Revenue increased by 15% compared to Q3", + "December was the strongest month", + "Product category A led growth" + ], + "charts_generated": 4, + "quality_score": 0.88 + }, + "tokens_used": { + "input_tokens": 3200, + "output_tokens": 1800, + "total_tokens": 5000 + }, + "cost_usd": 0.095, + "error_details": null, + "tools_used": ["data_analyzer"], + "retry_count": 0, + "metadata": { + "user_id": "user_456", + "session_id": "session_def", + "request_priority": "high" + } + }, + { + "task_id": "task_003", + "agent_id": "document_agent_1", + "task_type": "document_processing", + "task_description": "Extract key information from research paper PDF", + "start_time": "2024-01-15T09:10:00Z", + "end_time": "2024-01-15T09:12:20Z", + "duration_ms": 140000, + "status": "partial", + "actions": [ + { + "type": "tool_call", + "tool_name": "document_processor", + "duration_ms": 135000, + "success": true, + "parameters": { + "document_url": "https://example.com/research.pdf", + "processing_mode": "key_points" + } + }, + { + "type": "validation", + "description": "Validate extracted content", + "duration_ms": 5000, + "success": false, + "error": "Content validation failed - missing abstract" + } + ], + "results": { + "extracted_content": "Partial content extracted successfully", + "pages_processed": 12, + "validation_issues": ["Missing abstract section"], + "quality_score": 0.65 + }, + "tokens_used": { + "input_tokens": 5400, + "output_tokens": 3200, + "total_tokens": 8600 + }, + "cost_usd": 0.172, + "error_details": { + "error_type": "validation_error", + "error_message": "Document structure validation failed", + "affected_section": "abstract" + }, + "tools_used": ["document_processor"], + "retry_count": 1, + "metadata": { + "user_id": "user_789", + "session_id": "session_ghi", + "request_priority": "normal" + } + }, + { + "task_id": "task_004", + "agent_id": "communication_agent_1", + "task_type": "notification", + "task_description": "Send completion notification to project stakeholders", + "start_time": "2024-01-15T09:15:00Z", + "end_time": "2024-01-15T09:15:08Z", + "duration_ms": 8000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "notification_sender", + "duration_ms": 7500, + "success": true, + "parameters": { + "recipients": ["manager@example.com", "team@example.com"], + "message": "Project analysis completed successfully", + "channel": "email" + } + } + ], + "results": { + "notifications_sent": 2, + "delivery_confirmations": 2, + "quality_score": 1.0 + }, + "tokens_used": { + "input_tokens": 200, + "output_tokens": 150, + "total_tokens": 350 + }, + "cost_usd": 0.007, + "error_details": null, + "tools_used": ["notification_sender"], + "retry_count": 0, + "metadata": { + "user_id": "system", + "session_id": "session_jkl", + "request_priority": "normal" + } + }, + { + "task_id": "task_005", + "agent_id": "research_agent_2", + "task_type": "web_research", + "task_description": "Research competitive landscape analysis", + "start_time": "2024-01-15T09:20:00Z", + "end_time": "2024-01-15T09:25:30Z", + "duration_ms": 330000, + "status": "failure", + "actions": [ + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2800, + "success": true, + "parameters": { + "query": "competitive analysis software industry", + "limit": 15 + } + }, + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 30000, + "success": false, + "error": "Rate limit exceeded" + }, + { + "type": "retry", + "description": "Wait and retry search", + "duration_ms": 60000, + "success": false + }, + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 30000, + "success": false, + "error": "Service timeout" + } + ], + "results": { + "partial_results": "Initial search completed, subsequent searches failed", + "sources_found": 8, + "quality_score": 0.3 + }, + "tokens_used": { + "input_tokens": 800, + "output_tokens": 400, + "total_tokens": 1200 + }, + "cost_usd": 0.024, + "error_details": { + "error_type": "service_timeout", + "error_message": "Web search service exceeded timeout limit", + "retry_attempts": 2 + }, + "tools_used": ["web_search"], + "retry_count": 2, + "metadata": { + "user_id": "user_101", + "session_id": "session_mno", + "request_priority": "high" + } + }, + { + "task_id": "task_006", + "agent_id": "scheduler_agent_1", + "task_type": "task_scheduling", + "task_description": "Schedule weekly report generation", + "start_time": "2024-01-15T09:30:00Z", + "end_time": "2024-01-15T09:30:15Z", + "duration_ms": 15000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "task_scheduler", + "duration_ms": 12000, + "success": true, + "parameters": { + "task_definition": { + "action": "generate_report", + "parameters": {"report_type": "weekly_summary"} + }, + "schedule": { + "type": "recurring", + "recurrence_pattern": "weekly" + } + } + }, + { + "type": "validation", + "description": "Verify schedule creation", + "duration_ms": 3000, + "success": true + } + ], + "results": { + "task_scheduled": true, + "next_execution": "2024-01-22T09:30:00Z", + "schedule_id": "sched_789", + "quality_score": 1.0 + }, + "tokens_used": { + "input_tokens": 300, + "output_tokens": 200, + "total_tokens": 500 + }, + "cost_usd": 0.01, + "error_details": null, + "tools_used": ["task_scheduler"], + "retry_count": 0, + "metadata": { + "user_id": "user_202", + "session_id": "session_pqr", + "request_priority": "low" + } + }, + { + "task_id": "task_007", + "agent_id": "data_agent_2", + "task_type": "data_analysis", + "task_description": "Analyze customer satisfaction survey results", + "start_time": "2024-01-15T10:00:00Z", + "end_time": "2024-01-15T10:04:25Z", + "duration_ms": 265000, + "status": "timeout", + "actions": [ + { + "type": "data_ingestion", + "description": "Load survey response data", + "duration_ms": 15000, + "success": true + }, + { + "type": "tool_call", + "tool_name": "data_analyzer", + "duration_ms": 250000, + "success": false, + "error": "Operation timeout after 250 seconds" + } + ], + "results": { + "partial_analysis": "Data loaded but analysis incomplete", + "records_processed": 5000, + "total_records": 15000, + "quality_score": 0.2 + }, + "tokens_used": { + "input_tokens": 8000, + "output_tokens": 1000, + "total_tokens": 9000 + }, + "cost_usd": 0.18, + "error_details": { + "error_type": "timeout", + "error_message": "Data analysis operation exceeded maximum allowed time", + "timeout_limit_ms": 250000 + }, + "tools_used": ["data_analyzer"], + "retry_count": 0, + "metadata": { + "user_id": "user_303", + "session_id": "session_stu", + "request_priority": "normal" + } + }, + { + "task_id": "task_008", + "agent_id": "research_agent_1", + "task_type": "web_research", + "task_description": "Research industry best practices for remote work", + "start_time": "2024-01-15T10:30:00Z", + "end_time": "2024-01-15T10:33:15Z", + "duration_ms": 195000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2200, + "success": true, + "parameters": { + "query": "remote work best practices 2024", + "limit": 12 + } + }, + { + "type": "tool_call", + "tool_name": "web_search", + "duration_ms": 2400, + "success": true, + "parameters": { + "query": "hybrid work policies companies", + "limit": 8 + } + }, + { + "type": "content_synthesis", + "description": "Synthesize findings from multiple sources", + "duration_ms": 190400, + "success": true + } + ], + "results": { + "comprehensive_report": "Detailed analysis of remote work best practices with industry examples", + "sources_analyzed": 20, + "key_insights": 8, + "quality_score": 0.94 + }, + "tokens_used": { + "input_tokens": 2800, + "output_tokens": 4200, + "total_tokens": 7000 + }, + "cost_usd": 0.14, + "error_details": null, + "tools_used": ["web_search"], + "retry_count": 0, + "metadata": { + "user_id": "user_404", + "session_id": "session_vwx", + "request_priority": "normal" + } + }, + { + "task_id": "task_009", + "agent_id": "document_agent_2", + "task_type": "document_processing", + "task_description": "Process and summarize quarterly financial report", + "start_time": "2024-01-15T11:00:00Z", + "end_time": "2024-01-15T11:02:30Z", + "duration_ms": 150000, + "status": "success", + "actions": [ + { + "type": "tool_call", + "tool_name": "document_processor", + "duration_ms": 145000, + "success": true, + "parameters": { + "document_url": "https://example.com/q4-financial-report.pdf", + "processing_mode": "summary", + "output_format": "json" + } + }, + { + "type": "quality_check", + "description": "Validate summary completeness", + "duration_ms": 5000, + "success": true + } + ], + "results": { + "executive_summary": "Q4 revenue grew 12% YoY with strong performance in all segments", + "key_metrics_extracted": 15, + "summary_length": 500, + "quality_score": 0.91 + }, + "tokens_used": { + "input_tokens": 6500, + "output_tokens": 2200, + "total_tokens": 8700 + }, + "cost_usd": 0.174, + "error_details": null, + "tools_used": ["document_processor"], + "retry_count": 0, + "metadata": { + "user_id": "user_505", + "session_id": "session_yzA", + "request_priority": "high" + } + }, + { + "task_id": "task_010", + "agent_id": "communication_agent_2", + "task_type": "notification", + "task_description": "Send urgent system maintenance notification", + "start_time": "2024-01-15T11:30:00Z", + "end_time": "2024-01-15T11:30:45Z", + "duration_ms": 45000, + "status": "failure", + "actions": [ + { + "type": "tool_call", + "tool_name": "notification_sender", + "duration_ms": 30000, + "success": false, + "error": "Authentication failed - invalid API key", + "parameters": { + "recipients": ["all-users@example.com"], + "message": "Scheduled maintenance tonight 11 PM - 2 AM", + "channel": "email", + "priority": "urgent" + } + }, + { + "type": "retry", + "description": "Retry with backup credentials", + "duration_ms": 15000, + "success": false, + "error": "Backup authentication also failed" + } + ], + "results": { + "notifications_sent": 0, + "delivery_failures": 1, + "quality_score": 0.0 + }, + "tokens_used": { + "input_tokens": 150, + "output_tokens": 50, + "total_tokens": 200 + }, + "cost_usd": 0.004, + "error_details": { + "error_type": "authentication_error", + "error_message": "Failed to authenticate with notification service", + "retry_attempts": 1 + }, + "tools_used": ["notification_sender"], + "retry_count": 1, + "metadata": { + "user_id": "system", + "session_id": "session_BcD", + "request_priority": "urgent" + } + } + ] +} \ No newline at end of file diff --git a/skills/agent-designer/assets/sample_system_requirements.json b/skills/agent-designer/assets/sample_system_requirements.json new file mode 100644 index 00000000..0c14fcc2 --- /dev/null +++ b/skills/agent-designer/assets/sample_system_requirements.json @@ -0,0 +1,57 @@ +{ + "goal": "Build a comprehensive research and analysis platform that can gather information from multiple sources, analyze data, and generate detailed reports", + "description": "The system needs to handle complex research tasks involving web searches, data analysis, document processing, and collaborative report generation. It should be able to coordinate multiple specialists working in parallel while maintaining quality control and ensuring comprehensive coverage of research topics.", + "tasks": [ + "Conduct multi-source web research on specified topics", + "Analyze and synthesize information from various sources", + "Perform data processing and statistical analysis", + "Generate visualizations and charts from data", + "Create comprehensive written reports", + "Fact-check and validate information accuracy", + "Coordinate parallel research streams", + "Handle real-time information updates", + "Manage research project timelines", + "Provide interactive research assistance" + ], + "constraints": { + "max_response_time": 30000, + "budget_per_task": 1.0, + "quality_threshold": 0.9, + "concurrent_tasks": 10, + "data_retention_days": 90, + "security_level": "standard", + "compliance_requirements": ["GDPR", "data_minimization"] + }, + "team_size": 6, + "performance_requirements": { + "high_throughput": true, + "fault_tolerance": true, + "low_latency": false, + "scalability": "medium", + "availability": 0.99 + }, + "safety_requirements": [ + "Input validation and sanitization", + "Output content filtering", + "Rate limiting for external APIs", + "Error handling and graceful degradation", + "Human oversight for critical decisions", + "Audit logging for all operations" + ], + "integration_requirements": [ + "REST API endpoints for external systems", + "Webhook support for real-time updates", + "Database integration for data persistence", + "File storage for documents and media", + "Email notifications for important events", + "Dashboard for monitoring and control" + ], + "scale_requirements": { + "initial_users": 50, + "peak_concurrent_users": 200, + "data_volume_gb": 100, + "requests_per_hour": 1000, + "geographic_regions": ["US", "EU"], + "growth_projection": "50% per year" + } +} \ No newline at end of file diff --git a/skills/agent-designer/assets/sample_tool_descriptions.json b/skills/agent-designer/assets/sample_tool_descriptions.json new file mode 100644 index 00000000..ab055881 --- /dev/null +++ b/skills/agent-designer/assets/sample_tool_descriptions.json @@ -0,0 +1,545 @@ +{ + "tools": [ + { + "name": "web_search", + "purpose": "Search the web for information on specified topics with customizable filters and result limits", + "category": "search", + "inputs": [ + { + "name": "query", + "type": "string", + "description": "Search query string to find relevant information", + "required": true, + "min_length": 1, + "max_length": 500, + "examples": ["artificial intelligence trends", "climate change impact", "python programming tutorial"] + }, + { + "name": "limit", + "type": "integer", + "description": "Maximum number of search results to return", + "required": false, + "default": 10, + "minimum": 1, + "maximum": 100 + }, + { + "name": "language", + "type": "string", + "description": "Language code for search results", + "required": false, + "default": "en", + "enum": ["en", "es", "fr", "de", "it", "pt", "zh", "ja"] + }, + { + "name": "time_range", + "type": "string", + "description": "Time range filter for search results", + "required": false, + "enum": ["any", "day", "week", "month", "year"] + } + ], + "outputs": [ + { + "name": "results", + "type": "array", + "description": "Array of search result objects", + "items": { + "type": "object", + "properties": { + "title": {"type": "string"}, + "url": {"type": "string"}, + "snippet": {"type": "string"}, + "relevance_score": {"type": "number"} + } + } + }, + { + "name": "total_found", + "type": "integer", + "description": "Total number of results available" + } + ], + "error_conditions": [ + "Invalid query format", + "Network timeout", + "API rate limit exceeded", + "No results found", + "Service unavailable" + ], + "side_effects": [ + "Logs search query for analytics", + "May cache results temporarily" + ], + "idempotent": true, + "rate_limits": { + "requests_per_minute": 60, + "requests_per_hour": 1000, + "burst_limit": 10 + }, + "dependencies": [ + "search_api_service", + "content_filter_service" + ], + "examples": [ + { + "description": "Basic web search", + "input": { + "query": "machine learning algorithms", + "limit": 5 + }, + "expected_output": { + "results": [ + { + "title": "Introduction to Machine Learning Algorithms", + "url": "https://example.com/ml-intro", + "snippet": "Machine learning algorithms are computational methods...", + "relevance_score": 0.95 + } + ], + "total_found": 1250 + } + } + ], + "security_requirements": [ + "Query sanitization", + "Rate limiting by user", + "Content filtering" + ] + }, + { + "name": "data_analyzer", + "purpose": "Analyze structured data and generate statistical insights, trends, and visualizations", + "category": "data", + "inputs": [ + { + "name": "data", + "type": "object", + "description": "Structured data to analyze in JSON format", + "required": true, + "properties": { + "columns": {"type": "array"}, + "rows": {"type": "array"} + } + }, + { + "name": "analysis_type", + "type": "string", + "description": "Type of analysis to perform", + "required": true, + "enum": ["descriptive", "correlation", "trend", "distribution", "outlier_detection"] + }, + { + "name": "target_column", + "type": "string", + "description": "Primary column to focus analysis on", + "required": false + }, + { + "name": "include_visualization", + "type": "boolean", + "description": "Whether to generate visualization data", + "required": false, + "default": true + } + ], + "outputs": [ + { + "name": "insights", + "type": "array", + "description": "Array of analytical insights and findings" + }, + { + "name": "statistics", + "type": "object", + "description": "Statistical measures and metrics" + }, + { + "name": "visualization_data", + "type": "object", + "description": "Data formatted for visualization creation" + } + ], + "error_conditions": [ + "Invalid data format", + "Insufficient data points", + "Missing required columns", + "Data type mismatch", + "Analysis timeout" + ], + "side_effects": [ + "May create temporary analysis files", + "Logs analysis parameters for optimization" + ], + "idempotent": true, + "rate_limits": { + "requests_per_minute": 30, + "requests_per_hour": 500, + "burst_limit": 5 + }, + "dependencies": [ + "statistics_engine", + "visualization_service" + ], + "examples": [ + { + "description": "Basic descriptive analysis", + "input": { + "data": { + "columns": ["age", "salary", "department"], + "rows": [ + [25, 50000, "engineering"], + [30, 60000, "engineering"], + [28, 55000, "marketing"] + ] + }, + "analysis_type": "descriptive", + "target_column": "salary" + }, + "expected_output": { + "insights": [ + "Average salary is $55,000", + "Salary range: $50,000 - $60,000", + "Engineering department has higher average salary" + ], + "statistics": { + "mean": 55000, + "median": 55000, + "std_dev": 5000 + } + } + } + ], + "security_requirements": [ + "Data anonymization", + "Access control validation" + ] + }, + { + "name": "document_processor", + "purpose": "Process and extract information from various document formats including PDFs, Word docs, and plain text", + "category": "file", + "inputs": [ + { + "name": "document_url", + "type": "string", + "description": "URL or path to the document to process", + "required": true, + "pattern": "^(https?://|file://|/)" + }, + { + "name": "processing_mode", + "type": "string", + "description": "How to process the document", + "required": false, + "default": "full_text", + "enum": ["full_text", "summary", "key_points", "metadata_only"] + }, + { + "name": "output_format", + "type": "string", + "description": "Desired output format", + "required": false, + "default": "json", + "enum": ["json", "markdown", "plain_text"] + }, + { + "name": "language_detection", + "type": "boolean", + "description": "Whether to detect document language", + "required": false, + "default": true + } + ], + "outputs": [ + { + "name": "content", + "type": "string", + "description": "Extracted and processed document content" + }, + { + "name": "metadata", + "type": "object", + "description": "Document metadata including author, creation date, etc." + }, + { + "name": "language", + "type": "string", + "description": "Detected language of the document" + }, + { + "name": "word_count", + "type": "integer", + "description": "Total word count in the document" + } + ], + "error_conditions": [ + "Document not found", + "Unsupported file format", + "Document corrupted or unreadable", + "Access permission denied", + "Document too large" + ], + "side_effects": [ + "May download and cache documents temporarily", + "Creates processing logs for debugging" + ], + "idempotent": true, + "rate_limits": { + "requests_per_minute": 20, + "requests_per_hour": 300, + "burst_limit": 3 + }, + "dependencies": [ + "document_parser_service", + "language_detection_service", + "file_storage_service" + ], + "examples": [ + { + "description": "Process PDF document for full text extraction", + "input": { + "document_url": "https://example.com/research-paper.pdf", + "processing_mode": "full_text", + "output_format": "markdown" + }, + "expected_output": { + "content": "# Research Paper Title\n\nAbstract: This paper discusses...", + "metadata": { + "author": "Dr. Smith", + "creation_date": "2024-01-15", + "pages": 15 + }, + "language": "en", + "word_count": 3500 + } + } + ], + "security_requirements": [ + "URL validation", + "File type verification", + "Malware scanning", + "Access control enforcement" + ] + }, + { + "name": "notification_sender", + "purpose": "Send notifications via multiple channels including email, SMS, and webhooks", + "category": "communication", + "inputs": [ + { + "name": "recipients", + "type": "array", + "description": "List of recipient identifiers", + "required": true, + "min_items": 1, + "max_items": 100, + "items": { + "type": "string", + "pattern": "^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$|^\\+?[1-9]\\d{1,14}$" + } + }, + { + "name": "message", + "type": "string", + "description": "Message content to send", + "required": true, + "min_length": 1, + "max_length": 10000 + }, + { + "name": "channel", + "type": "string", + "description": "Communication channel to use", + "required": false, + "default": "email", + "enum": ["email", "sms", "webhook", "push"] + }, + { + "name": "priority", + "type": "string", + "description": "Message priority level", + "required": false, + "default": "normal", + "enum": ["low", "normal", "high", "urgent"] + }, + { + "name": "template_id", + "type": "string", + "description": "Optional template ID for formatting", + "required": false + } + ], + "outputs": [ + { + "name": "delivery_status", + "type": "object", + "description": "Status of message delivery to each recipient" + }, + { + "name": "message_id", + "type": "string", + "description": "Unique identifier for the sent message" + }, + { + "name": "delivery_timestamp", + "type": "string", + "description": "ISO timestamp when message was sent" + } + ], + "error_conditions": [ + "Invalid recipient format", + "Message too long", + "Channel service unavailable", + "Authentication failure", + "Rate limit exceeded for channel" + ], + "side_effects": [ + "Sends actual notifications to recipients", + "Logs delivery attempts and results", + "Updates delivery statistics" + ], + "idempotent": false, + "rate_limits": { + "requests_per_minute": 100, + "requests_per_hour": 2000, + "burst_limit": 20 + }, + "dependencies": [ + "email_service", + "sms_service", + "webhook_service" + ], + "examples": [ + { + "description": "Send email notification", + "input": { + "recipients": ["user@example.com"], + "message": "Your report has been completed and is ready for review.", + "channel": "email", + "priority": "normal" + }, + "expected_output": { + "delivery_status": { + "user@example.com": "delivered" + }, + "message_id": "msg_12345", + "delivery_timestamp": "2024-01-15T10:30:00Z" + } + } + ], + "security_requirements": [ + "Recipient validation", + "Message content filtering", + "Rate limiting per user", + "Delivery confirmation" + ] + }, + { + "name": "task_scheduler", + "purpose": "Schedule and manage delayed or recurring tasks within the agent system", + "category": "compute", + "inputs": [ + { + "name": "task_definition", + "type": "object", + "description": "Definition of the task to be scheduled", + "required": true, + "properties": { + "action": {"type": "string"}, + "parameters": {"type": "object"}, + "retry_policy": {"type": "object"} + } + }, + { + "name": "schedule", + "type": "object", + "description": "Scheduling parameters for the task", + "required": true, + "properties": { + "type": {"type": "string", "enum": ["once", "recurring"]}, + "execute_at": {"type": "string"}, + "recurrence_pattern": {"type": "string"} + } + }, + { + "name": "priority", + "type": "integer", + "description": "Task priority (1-10, higher is more urgent)", + "required": false, + "default": 5, + "minimum": 1, + "maximum": 10 + } + ], + "outputs": [ + { + "name": "task_id", + "type": "string", + "description": "Unique identifier for the scheduled task" + }, + { + "name": "next_execution", + "type": "string", + "description": "ISO timestamp of next scheduled execution" + }, + { + "name": "status", + "type": "string", + "description": "Current status of the scheduled task" + } + ], + "error_conditions": [ + "Invalid schedule format", + "Past execution time specified", + "Task queue full", + "Invalid task definition", + "Scheduling service unavailable" + ], + "side_effects": [ + "Creates scheduled tasks in the system", + "May consume system resources for task storage", + "Updates scheduling metrics" + ], + "idempotent": false, + "rate_limits": { + "requests_per_minute": 50, + "requests_per_hour": 1000, + "burst_limit": 10 + }, + "dependencies": [ + "task_scheduler_service", + "task_executor_service" + ], + "examples": [ + { + "description": "Schedule a one-time report generation", + "input": { + "task_definition": { + "action": "generate_report", + "parameters": { + "report_type": "monthly_summary", + "recipients": ["manager@example.com"] + } + }, + "schedule": { + "type": "once", + "execute_at": "2024-02-01T09:00:00Z" + }, + "priority": 7 + }, + "expected_output": { + "task_id": "task_67890", + "next_execution": "2024-02-01T09:00:00Z", + "status": "scheduled" + } + } + ], + "security_requirements": [ + "Task definition validation", + "User authorization for scheduling", + "Resource limit enforcement" + ] + } + ] +} \ No newline at end of file diff --git a/skills/agent-designer/expected_outputs/sample_agent_architecture.json b/skills/agent-designer/expected_outputs/sample_agent_architecture.json new file mode 100644 index 00000000..0af7c666 --- /dev/null +++ b/skills/agent-designer/expected_outputs/sample_agent_architecture.json @@ -0,0 +1,488 @@ +{ + "architecture_design": { + "pattern": "supervisor", + "agents": [ + { + "name": "supervisor_agent", + "role": "Task Coordinator and Quality Controller", + "archetype": "coordinator", + "responsibilities": [ + "task_decomposition", + "delegation", + "progress_monitoring", + "quality_assurance", + "result_aggregation" + ], + "capabilities": [ + "planning", + "coordination", + "evaluation", + "decision_making" + ], + "tools": [ + { + "name": "file_manager", + "description": "Manage files and directories", + "input_schema": { + "type": "object", + "properties": { + "action": { + "type": "string" + }, + "path": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "success": { + "type": "boolean" + }, + "content": { + "type": "string" + } + } + }, + "capabilities": [ + "file_operations", + "data_management" + ], + "reliability": "high", + "latency": "low" + }, + { + "name": "data_analyzer", + "description": "Analyze and process data", + "input_schema": { + "type": "object", + "properties": { + "data": { + "type": "object" + }, + "analysis_type": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "insights": { + "type": "array" + }, + "metrics": { + "type": "object" + } + } + }, + "capabilities": [ + "data_analysis", + "statistics", + "visualization" + ], + "reliability": "high", + "latency": "medium" + } + ], + "communication_interfaces": [ + "user_interface", + "agent_messaging" + ], + "constraints": { + "max_concurrent_supervisions": 5, + "decision_timeout": "30s" + }, + "success_criteria": [ + "successful task completion", + "optimal resource utilization", + "quality standards met" + ], + "dependencies": [] + }, + { + "name": "research_specialist", + "role": "Research Specialist", + "archetype": "specialist", + "responsibilities": [ + "Conduct multi-source web research on specified topics", + "Handle real-time information updates" + ], + "capabilities": [ + "research_expertise", + "specialized_tools", + "domain_knowledge" + ], + "tools": [ + { + "name": "web_search", + "description": "Search the web for information", + "input_schema": { + "type": "object", + "properties": { + "query": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "results": { + "type": "array" + } + } + }, + "capabilities": [ + "research", + "information_gathering" + ], + "reliability": "high", + "latency": "medium" + }, + { + "name": "data_analyzer", + "description": "Analyze and process data", + "input_schema": { + "type": "object", + "properties": { + "data": { + "type": "object" + }, + "analysis_type": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "insights": { + "type": "array" + }, + "metrics": { + "type": "object" + } + } + }, + "capabilities": [ + "data_analysis", + "statistics", + "visualization" + ], + "reliability": "high", + "latency": "medium" + } + ], + "communication_interfaces": [ + "supervisor_messaging" + ], + "constraints": { + "domain_scope": "research", + "task_queue_size": 10 + }, + "success_criteria": [ + "excel in research tasks", + "maintain domain expertise", + "provide quality output" + ], + "dependencies": [ + "supervisor_agent" + ] + }, + { + "name": "data_specialist", + "role": "Data Specialist", + "archetype": "specialist", + "responsibilities": [ + "Analyze and synthesize information from various sources", + "Perform data processing and statistical analysis", + "Generate visualizations and charts from data" + ], + "capabilities": [ + "data_expertise", + "specialized_tools", + "domain_knowledge" + ], + "tools": [ + { + "name": "data_analyzer", + "description": "Analyze and process data", + "input_schema": { + "type": "object", + "properties": { + "data": { + "type": "object" + }, + "analysis_type": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "insights": { + "type": "array" + }, + "metrics": { + "type": "object" + } + } + }, + "capabilities": [ + "data_analysis", + "statistics", + "visualization" + ], + "reliability": "high", + "latency": "medium" + }, + { + "name": "file_manager", + "description": "Manage files and directories", + "input_schema": { + "type": "object", + "properties": { + "action": { + "type": "string" + }, + "path": { + "type": "string" + } + } + }, + "output_schema": { + "type": "object", + "properties": { + "success": { + "type": "boolean" + }, + "content": { + "type": "string" + } + } + }, + "capabilities": [ + "file_operations", + "data_management" + ], + "reliability": "high", + "latency": "low" + } + ], + "communication_interfaces": [ + "supervisor_messaging" + ], + "constraints": { + "domain_scope": "data", + "task_queue_size": 10 + }, + "success_criteria": [ + "excel in data tasks", + "maintain domain expertise", + "provide quality output" + ], + "dependencies": [ + "supervisor_agent" + ] + } + ], + "communication_topology": [ + { + "from_agent": "supervisor_agent", + "to_agent": "research_specialist", + "pattern": "direct_message", + "data_format": "json", + "frequency": "on_demand", + "criticality": "high" + }, + { + "from_agent": "research_specialist", + "to_agent": "supervisor_agent", + "pattern": "direct_message", + "data_format": "json", + "frequency": "on_completion", + "criticality": "high" + }, + { + "from_agent": "supervisor_agent", + "to_agent": "data_specialist", + "pattern": "direct_message", + "data_format": "json", + "frequency": "on_demand", + "criticality": "high" + }, + { + "from_agent": "data_specialist", + "to_agent": "supervisor_agent", + "pattern": "direct_message", + "data_format": "json", + "frequency": "on_completion", + "criticality": "high" + } + ], + "shared_resources": [ + { + "type": "message_queue", + "capacity": 1000 + }, + { + "type": "shared_memory", + "size": "1GB" + }, + { + "type": "event_store", + "retention": "30 days" + } + ], + "guardrails": [ + { + "type": "input_validation", + "rules": "strict_schema_enforcement" + }, + { + "type": "rate_limiting", + "limit": "100_requests_per_minute" + }, + { + "type": "output_filtering", + "rules": "content_safety_check" + } + ], + "scaling_strategy": { + "horizontal_scaling": true, + "auto_scaling_triggers": [ + "cpu > 80%", + "queue_depth > 100" + ], + "max_instances_per_agent": 5 + }, + "failure_handling": { + "retry_policy": "exponential_backoff", + "circuit_breaker": true, + "fallback_strategies": [ + "graceful_degradation", + "human_escalation" + ] + } + }, + "mermaid_diagram": "graph TD\n supervisor_agent[Task Coordinator and Quality Controller]:::coordinator\n research_specialist[Research Specialist]:::specialist\n data_specialist[Data Specialist]:::specialist\n supervisor_agent --> research_specialist\n research_specialist --> supervisor_agent\n supervisor_agent --> data_specialist\n data_specialist --> supervisor_agent\n\n classDef coordinator fill:#e1f5fe,stroke:#01579b,stroke-width:2px\n classDef specialist fill:#f3e5f5,stroke:#4a148c,stroke-width:2px\n classDef interface fill:#e8f5e8,stroke:#1b5e20,stroke-width:2px\n classDef monitor fill:#fff3e0,stroke:#e65100,stroke-width:2px\n class supervisor_agent coordinator\n class research_specialist specialist\n class data_specialist specialist", + "implementation_roadmap": { + "total_duration": "8-12 weeks", + "phases": [ + { + "phase": 1, + "name": "Core Infrastructure", + "duration": "2-3 weeks", + "tasks": [ + "Set up development environment", + "Implement basic agent framework", + "Create communication infrastructure", + "Set up monitoring and logging", + "Implement basic tools" + ], + "deliverables": [ + "Agent runtime framework", + "Communication layer", + "Basic monitoring dashboard" + ] + }, + { + "phase": 2, + "name": "Agent Implementation", + "duration": "3-4 weeks", + "tasks": [ + "Implement individual agent logic", + "Create agent-specific tools", + "Implement communication protocols", + "Add error handling and recovery", + "Create agent configuration system" + ], + "deliverables": [ + "Functional agent implementations", + "Tool integration", + "Configuration management" + ] + }, + { + "phase": 3, + "name": "Integration and Testing", + "duration": "2-3 weeks", + "tasks": [ + "Integrate all agents", + "End-to-end testing", + "Performance optimization", + "Security implementation", + "Documentation creation" + ], + "deliverables": [ + "Integrated system", + "Test suite", + "Performance benchmarks", + "Security audit report" + ] + }, + { + "phase": 4, + "name": "Deployment and Monitoring", + "duration": "1-2 weeks", + "tasks": [ + "Production deployment", + "Monitoring setup", + "Alerting configuration", + "User training", + "Go-live support" + ], + "deliverables": [ + "Production system", + "Monitoring dashboard", + "Operational runbooks", + "Training materials" + ] + } + ], + "critical_path": [ + "Agent framework implementation", + "Communication layer development", + "Integration testing", + "Production deployment" + ], + "risks": [ + { + "risk": "Communication complexity", + "impact": "high", + "mitigation": "Start with simple protocols, iterate" + }, + { + "risk": "Agent coordination failures", + "impact": "medium", + "mitigation": "Implement robust error handling and fallbacks" + }, + { + "risk": "Performance bottlenecks", + "impact": "medium", + "mitigation": "Early performance testing and optimization" + } + ], + "success_criteria": [ + "Input validation and sanitization", + "Output content filtering", + "Rate limiting for external APIs", + "Error handling and graceful degradation", + "Human oversight for critical decisions", + "Audit logging for all operations", + "All agents operational", + "Communication working reliably", + "Performance targets met", + "Error rate below 1%" + ] + }, + "metadata": { + "generated_by": "agent_planner.py", + "requirements_file": "sample_system_requirements.json", + "architecture_pattern": "supervisor", + "agent_count": 3 + } +} \ No newline at end of file diff --git a/skills/agent-designer/expected_outputs/sample_evaluation_report.json b/skills/agent-designer/expected_outputs/sample_evaluation_report.json new file mode 100644 index 00000000..0c9bce7f --- /dev/null +++ b/skills/agent-designer/expected_outputs/sample_evaluation_report.json @@ -0,0 +1,570 @@ +{ + "summary": { + "evaluation_period": { + "start_time": "2024-01-15T09:00:00Z", + "end_time": "2024-01-15T11:30:45Z", + "total_duration_hours": 2.51 + }, + "overall_health": "good", + "key_findings": [ + "Success rate (80.0%) below target", + "High average latency (16.9s)", + "2 high-impact error patterns identified" + ], + "critical_issues": 0, + "improvement_opportunities": 6 + }, + "system_metrics": { + "total_tasks": 10, + "successful_tasks": 8, + "failed_tasks": 2, + "partial_tasks": 1, + "timeout_tasks": 1, + "success_rate": 0.8, + "failure_rate": 0.2, + "average_duration_ms": 169800.0, + "median_duration_ms": 152500.0, + "percentile_95_duration_ms": 330000.0, + "min_duration_ms": 8000, + "max_duration_ms": 330000, + "total_tokens_used": 53700, + "average_tokens_per_task": 5370.0, + "total_cost_usd": 1.074, + "average_cost_per_task": 0.1074, + "cost_per_token": 0.00002, + "throughput_tasks_per_hour": 3.98, + "error_rate": 0.3, + "retry_rate": 0.3 + }, + "agent_metrics": { + "research_agent_1": { + "total_tasks": 2, + "successful_tasks": 2, + "failed_tasks": 0, + "partial_tasks": 0, + "timeout_tasks": 0, + "success_rate": 1.0, + "failure_rate": 0.0, + "average_duration_ms": 174500.0, + "median_duration_ms": 174500.0, + "percentile_95_duration_ms": 195000.0, + "min_duration_ms": 154000, + "max_duration_ms": 195000, + "total_tokens_used": 11050, + "average_tokens_per_task": 5525.0, + "total_cost_usd": 0.221, + "average_cost_per_task": 0.1105, + "cost_per_token": 0.00002, + "throughput_tasks_per_hour": 11.49, + "error_rate": 0.0, + "retry_rate": 0.0 + }, + "data_agent_1": { + "total_tasks": 1, + "successful_tasks": 1, + "failed_tasks": 0, + "partial_tasks": 0, + "timeout_tasks": 0, + "success_rate": 1.0, + "failure_rate": 0.0, + "average_duration_ms": 165000.0, + "median_duration_ms": 165000.0, + "percentile_95_duration_ms": 165000.0, + "min_duration_ms": 165000, + "max_duration_ms": 165000, + "total_tokens_used": 5000, + "average_tokens_per_task": 5000.0, + "total_cost_usd": 0.095, + "average_cost_per_task": 0.095, + "cost_per_token": 0.000019, + "throughput_tasks_per_hour": 21.82, + "error_rate": 0.0, + "retry_rate": 0.0 + }, + "document_agent_1": { + "total_tasks": 1, + "successful_tasks": 0, + "failed_tasks": 0, + "partial_tasks": 1, + "timeout_tasks": 0, + "success_rate": 0.0, + "failure_rate": 0.0, + "average_duration_ms": 140000.0, + "median_duration_ms": 140000.0, + "percentile_95_duration_ms": 140000.0, + "min_duration_ms": 140000, + "max_duration_ms": 140000, + "total_tokens_used": 8600, + "average_tokens_per_task": 8600.0, + "total_cost_usd": 0.172, + "average_cost_per_task": 0.172, + "cost_per_token": 0.00002, + "throughput_tasks_per_hour": 25.71, + "error_rate": 1.0, + "retry_rate": 1.0 + } + }, + "task_type_metrics": { + "web_research": { + "total_tasks": 3, + "successful_tasks": 2, + "failed_tasks": 1, + "partial_tasks": 0, + "timeout_tasks": 0, + "success_rate": 0.667, + "failure_rate": 0.333, + "average_duration_ms": 226333.33, + "median_duration_ms": 195000.0, + "percentile_95_duration_ms": 330000.0, + "min_duration_ms": 154000, + "max_duration_ms": 330000, + "total_tokens_used": 12250, + "average_tokens_per_task": 4083.33, + "total_cost_usd": 0.245, + "average_cost_per_task": 0.082, + "cost_per_token": 0.00002, + "throughput_tasks_per_hour": 2.65, + "error_rate": 0.333, + "retry_rate": 0.333 + }, + "data_analysis": { + "total_tasks": 2, + "successful_tasks": 1, + "failed_tasks": 0, + "partial_tasks": 0, + "timeout_tasks": 1, + "success_rate": 0.5, + "failure_rate": 0.0, + "average_duration_ms": 215000.0, + "median_duration_ms": 215000.0, + "percentile_95_duration_ms": 265000.0, + "min_duration_ms": 165000, + "max_duration_ms": 265000, + "total_tokens_used": 14000, + "average_tokens_per_task": 7000.0, + "total_cost_usd": 0.275, + "average_cost_per_task": 0.138, + "cost_per_token": 0.0000196, + "throughput_tasks_per_hour": 1.86, + "error_rate": 0.5, + "retry_rate": 0.0 + } + }, + "tool_usage_analysis": { + "web_search": { + "usage_count": 3, + "error_rate": 0.333, + "avg_duration": 126666.67, + "affected_workflows": [ + "web_research" + ], + "retry_count": 2 + }, + "data_analyzer": { + "usage_count": 2, + "error_rate": 0.0, + "avg_duration": 205000.0, + "affected_workflows": [ + "data_analysis" + ], + "retry_count": 0 + }, + "document_processor": { + "usage_count": 2, + "error_rate": 0.0, + "avg_duration": 140000.0, + "affected_workflows": [ + "document_processing" + ], + "retry_count": 1 + }, + "notification_sender": { + "usage_count": 2, + "error_rate": 0.5, + "avg_duration": 18750.0, + "affected_workflows": [ + "notification" + ], + "retry_count": 1 + }, + "task_scheduler": { + "usage_count": 1, + "error_rate": 0.0, + "avg_duration": 12000.0, + "affected_workflows": [ + "task_scheduling" + ], + "retry_count": 0 + } + }, + "error_analysis": [ + { + "error_type": "timeout", + "count": 2, + "percentage": 20.0, + "affected_agents": [ + "research_agent_2", + "data_agent_2" + ], + "affected_task_types": [ + "web_research", + "data_analysis" + ], + "common_patterns": [ + "timeout", + "exceeded", + "limit" + ], + "suggested_fixes": [ + "Increase timeout values", + "Optimize slow operations", + "Add retry logic with exponential backoff", + "Parallelize independent operations" + ], + "impact_level": "high" + }, + { + "error_type": "authentication", + "count": 1, + "percentage": 10.0, + "affected_agents": [ + "communication_agent_2" + ], + "affected_task_types": [ + "notification" + ], + "common_patterns": [ + "authentication", + "failed", + "invalid" + ], + "suggested_fixes": [ + "Check credential rotation", + "Implement token refresh logic", + "Add authentication retry", + "Verify permission scopes" + ], + "impact_level": "high" + }, + { + "error_type": "validation", + "count": 1, + "percentage": 10.0, + "affected_agents": [ + "document_agent_1" + ], + "affected_task_types": [ + "document_processing" + ], + "common_patterns": [ + "validation", + "failed", + "missing" + ], + "suggested_fixes": [ + "Strengthen input validation", + "Add data sanitization", + "Improve error messages", + "Add input examples" + ], + "impact_level": "medium" + } + ], + "bottleneck_analysis": [ + { + "bottleneck_type": "tool", + "location": "notification_sender", + "severity": "medium", + "description": "Tool notification_sender has high error rate (50.0%)", + "impact_on_performance": { + "reliability_impact": 1.0, + "retry_overhead": 1000 + }, + "affected_workflows": [ + "notification" + ], + "optimization_suggestions": [ + "Review tool implementation", + "Add better error handling for tool", + "Implement tool fallbacks", + "Consider alternative tools" + ], + "estimated_improvement": { + "error_reduction": 0.35, + "performance_gain": 1.2 + } + }, + { + "bottleneck_type": "tool", + "location": "web_search", + "severity": "medium", + "description": "Tool web_search has high error rate (33.3%)", + "impact_on_performance": { + "reliability_impact": 1.0, + "retry_overhead": 2000 + }, + "affected_workflows": [ + "web_research" + ], + "optimization_suggestions": [ + "Review tool implementation", + "Add better error handling for tool", + "Implement tool fallbacks", + "Consider alternative tools" + ], + "estimated_improvement": { + "error_reduction": 0.233, + "performance_gain": 1.2 + } + } + ], + "optimization_recommendations": [ + { + "category": "reliability", + "priority": "high", + "title": "Improve System Reliability", + "description": "System success rate is 80.0%, below target of 90%", + "implementation_effort": "medium", + "expected_impact": { + "success_rate_improvement": 0.1, + "cost_reduction": 0.01611 + }, + "estimated_cost_savings": 0.1074, + "estimated_performance_gain": 1.2, + "implementation_steps": [ + "Identify and fix top error patterns", + "Implement better error handling and retries", + "Add comprehensive monitoring and alerting", + "Implement graceful degradation patterns" + ], + "risks": [ + "Temporary increase in complexity", + "Potential initial performance overhead" + ], + "prerequisites": [ + "Error analysis completion", + "Monitoring infrastructure" + ] + }, + { + "category": "performance", + "priority": "high", + "title": "Reduce Task Latency", + "description": "Average task duration (169.8s) exceeds target", + "implementation_effort": "high", + "expected_impact": { + "latency_reduction": 0.49, + "throughput_improvement": 1.5 + }, + "estimated_performance_gain": 1.4, + "implementation_steps": [ + "Profile and optimize slow operations", + "Implement parallel processing where possible", + "Add caching for expensive operations", + "Optimize API calls and reduce round trips" + ], + "risks": [ + "Increased system complexity", + "Potential resource usage increase" + ], + "prerequisites": [ + "Performance profiling tools", + "Caching infrastructure" + ] + }, + { + "category": "cost", + "priority": "medium", + "title": "Optimize Token Usage and Costs", + "description": "Average cost per task ($0.107) is above optimal range", + "implementation_effort": "low", + "expected_impact": { + "cost_reduction": 0.032, + "efficiency_improvement": 1.15 + }, + "estimated_cost_savings": 0.322, + "estimated_performance_gain": 1.05, + "implementation_steps": [ + "Implement prompt optimization", + "Add response caching for repeated queries", + "Use smaller models for simple tasks", + "Implement token usage monitoring and alerts" + ], + "risks": [ + "Potential quality reduction with smaller models" + ], + "prerequisites": [ + "Token usage analysis", + "Caching infrastructure" + ] + }, + { + "category": "reliability", + "priority": "high", + "title": "Address Timeout Errors", + "description": "Timeout errors occur in 20.0% of cases", + "implementation_effort": "medium", + "expected_impact": { + "error_reduction": 0.2, + "reliability_improvement": 1.1 + }, + "estimated_cost_savings": 0.1074, + "implementation_steps": [ + "Increase timeout values", + "Optimize slow operations", + "Add retry logic with exponential backoff", + "Parallelize independent operations" + ], + "risks": [ + "May require significant code changes" + ], + "prerequisites": [ + "Root cause analysis", + "Testing framework" + ] + }, + { + "category": "reliability", + "priority": "high", + "title": "Address Authentication Errors", + "description": "Authentication errors occur in 10.0% of cases", + "implementation_effort": "medium", + "expected_impact": { + "error_reduction": 0.1, + "reliability_improvement": 1.1 + }, + "estimated_cost_savings": 0.1074, + "implementation_steps": [ + "Check credential rotation", + "Implement token refresh logic", + "Add authentication retry", + "Verify permission scopes" + ], + "risks": [ + "May require significant code changes" + ], + "prerequisites": [ + "Root cause analysis", + "Testing framework" + ] + }, + { + "category": "performance", + "priority": "medium", + "title": "Address Tool Bottleneck", + "description": "Tool notification_sender has high error rate (50.0%)", + "implementation_effort": "medium", + "expected_impact": { + "error_reduction": 0.35, + "performance_gain": 1.2 + }, + "estimated_performance_gain": 1.2, + "implementation_steps": [ + "Review tool implementation", + "Add better error handling for tool", + "Implement tool fallbacks", + "Consider alternative tools" + ], + "risks": [ + "System downtime during implementation", + "Potential cascade effects" + ], + "prerequisites": [ + "Impact assessment", + "Rollback plan" + ] + } + ], + "trends_analysis": { + "daily_success_rates": { + "2024-01-15": 0.8 + }, + "daily_avg_durations": { + "2024-01-15": 169800.0 + }, + "daily_costs": { + "2024-01-15": 1.074 + }, + "trend_direction": { + "success_rate": "stable", + "duration": "stable", + "cost": "stable" + } + }, + "cost_breakdown": { + "total_cost": 1.074, + "cost_by_agent": { + "research_agent_1": 0.221, + "research_agent_2": 0.024, + "data_agent_1": 0.095, + "data_agent_2": 0.18, + "document_agent_1": 0.172, + "document_agent_2": 0.174, + "communication_agent_1": 0.007, + "communication_agent_2": 0.004, + "scheduler_agent_1": 0.01 + }, + "cost_by_task_type": { + "web_research": 0.245, + "data_analysis": 0.275, + "document_processing": 0.346, + "notification": 0.011, + "task_scheduling": 0.01 + }, + "cost_per_token": 0.00002, + "top_cost_drivers": [ + [ + "document_processing", + 0.346 + ], + [ + "data_analysis", + 0.275 + ], + [ + "web_research", + 0.245 + ], + [ + "notification", + 0.011 + ], + [ + "task_scheduling", + 0.01 + ] + ] + }, + "sla_compliance": { + "overall_compliant": false, + "sla_details": { + "success_rate": { + "target": 0.95, + "actual": 0.8, + "compliant": false, + "gap": 0.15 + }, + "average_latency": { + "target": 10000, + "actual": 169800.0, + "compliant": false, + "gap": 159800.0 + }, + "error_rate": { + "target": 0.05, + "actual": 0.3, + "compliant": false, + "gap": 0.25 + } + }, + "compliance_score": 0.0 + }, + "metadata": { + "generated_at": "2024-01-15T12:00:00Z", + "evaluator_version": "1.0", + "total_logs_processed": 10, + "agents_analyzed": 9, + "task_types_analyzed": 5, + "analysis_completeness": "full" + } +} \ No newline at end of file diff --git a/skills/agent-designer/expected_outputs/sample_tool_schemas.json b/skills/agent-designer/expected_outputs/sample_tool_schemas.json new file mode 100644 index 00000000..72175c79 --- /dev/null +++ b/skills/agent-designer/expected_outputs/sample_tool_schemas.json @@ -0,0 +1,416 @@ +{ + "tool_schemas": [ + { + "name": "web_search", + "description": "Search the web for information on specified topics with customizable filters and result limits", + "openai_schema": { + "name": "web_search", + "description": "Search the web for information on specified topics with customizable filters and result limits", + "parameters": { + "type": "object", + "properties": { + "query": { + "type": "string", + "description": "Search query string to find relevant information", + "minLength": 1, + "maxLength": 500, + "examples": [ + "artificial intelligence trends", + "climate change impact", + "python programming tutorial" + ] + }, + "limit": { + "type": "integer", + "description": "Maximum number of search results to return", + "minimum": 1, + "maximum": 100, + "default": 10 + }, + "language": { + "type": "string", + "description": "Language code for search results", + "enum": [ + "en", + "es", + "fr", + "de", + "it", + "pt", + "zh", + "ja" + ], + "default": "en" + }, + "time_range": { + "type": "string", + "description": "Time range filter for search results", + "enum": [ + "any", + "day", + "week", + "month", + "year" + ] + } + }, + "required": [ + "query" + ], + "additionalProperties": false + } + }, + "anthropic_schema": { + "name": "web_search", + "description": "Search the web for information on specified topics with customizable filters and result limits", + "input_schema": { + "type": "object", + "properties": { + "query": { + "type": "string", + "description": "Search query string to find relevant information", + "minLength": 1, + "maxLength": 500 + }, + "limit": { + "type": "integer", + "description": "Maximum number of search results to return", + "minimum": 1, + "maximum": 100 + }, + "language": { + "type": "string", + "description": "Language code for search results", + "enum": [ + "en", + "es", + "fr", + "de", + "it", + "pt", + "zh", + "ja" + ] + }, + "time_range": { + "type": "string", + "description": "Time range filter for search results", + "enum": [ + "any", + "day", + "week", + "month", + "year" + ] + } + }, + "required": [ + "query" + ] + } + }, + "validation_rules": [ + { + "parameter": "query", + "rules": { + "minLength": 1, + "maxLength": 500 + } + }, + { + "parameter": "limit", + "rules": { + "minimum": 1, + "maximum": 100 + } + } + ], + "error_responses": [ + { + "error_code": "invalid_input", + "error_message": "Invalid input parameters provided", + "http_status": 400, + "retry_after": null, + "details": { + "validation_errors": [] + } + }, + { + "error_code": "authentication_required", + "error_message": "Authentication required to access this tool", + "http_status": 401, + "retry_after": null, + "details": null + }, + { + "error_code": "rate_limit_exceeded", + "error_message": "Rate limit exceeded. Please try again later", + "http_status": 429, + "retry_after": 60, + "details": null + } + ], + "rate_limits": { + "requests_per_minute": 60, + "requests_per_hour": 1000, + "requests_per_day": 10000, + "burst_limit": 10, + "cooldown_period": 60, + "rate_limit_key": "user_id" + }, + "examples": [ + { + "description": "Basic web search", + "input": { + "query": "machine learning algorithms", + "limit": 5 + }, + "expected_output": { + "results": [ + { + "title": "Introduction to Machine Learning Algorithms", + "url": "https://example.com/ml-intro", + "snippet": "Machine learning algorithms are computational methods...", + "relevance_score": 0.95 + } + ], + "total_found": 1250 + } + } + ], + "metadata": { + "category": "search", + "idempotent": true, + "side_effects": [ + "Logs search query for analytics", + "May cache results temporarily" + ], + "dependencies": [ + "search_api_service", + "content_filter_service" + ], + "security_requirements": [ + "Query sanitization", + "Rate limiting by user", + "Content filtering" + ], + "generated_at": "2024-01-15T10:30:00Z", + "schema_version": "1.0", + "input_parameters": 4, + "output_parameters": 2, + "required_parameters": 1, + "optional_parameters": 3 + } + }, + { + "name": "data_analyzer", + "description": "Analyze structured data and generate statistical insights, trends, and visualizations", + "openai_schema": { + "name": "data_analyzer", + "description": "Analyze structured data and generate statistical insights, trends, and visualizations", + "parameters": { + "type": "object", + "properties": { + "data": { + "type": "object", + "description": "Structured data to analyze in JSON format", + "properties": { + "columns": { + "type": "array" + }, + "rows": { + "type": "array" + } + }, + "additionalProperties": false + }, + "analysis_type": { + "type": "string", + "description": "Type of analysis to perform", + "enum": [ + "descriptive", + "correlation", + "trend", + "distribution", + "outlier_detection" + ] + }, + "target_column": { + "type": "string", + "description": "Primary column to focus analysis on", + "maxLength": 1000 + }, + "include_visualization": { + "type": "boolean", + "description": "Whether to generate visualization data", + "default": true + } + }, + "required": [ + "data", + "analysis_type" + ], + "additionalProperties": false + } + }, + "anthropic_schema": { + "name": "data_analyzer", + "description": "Analyze structured data and generate statistical insights, trends, and visualizations", + "input_schema": { + "type": "object", + "properties": { + "data": { + "type": "object", + "description": "Structured data to analyze in JSON format" + }, + "analysis_type": { + "type": "string", + "description": "Type of analysis to perform", + "enum": [ + "descriptive", + "correlation", + "trend", + "distribution", + "outlier_detection" + ] + }, + "target_column": { + "type": "string", + "description": "Primary column to focus analysis on", + "maxLength": 1000 + }, + "include_visualization": { + "type": "boolean", + "description": "Whether to generate visualization data" + } + }, + "required": [ + "data", + "analysis_type" + ] + } + }, + "validation_rules": [ + { + "parameter": "target_column", + "rules": { + "maxLength": 1000 + } + } + ], + "error_responses": [ + { + "error_code": "invalid_input", + "error_message": "Invalid input parameters provided", + "http_status": 400, + "retry_after": null, + "details": { + "validation_errors": [] + } + }, + { + "error_code": "authentication_required", + "error_message": "Authentication required to access this tool", + "http_status": 401, + "retry_after": null, + "details": null + }, + { + "error_code": "rate_limit_exceeded", + "error_message": "Rate limit exceeded. Please try again later", + "http_status": 429, + "retry_after": 60, + "details": null + } + ], + "rate_limits": { + "requests_per_minute": 30, + "requests_per_hour": 500, + "requests_per_day": 5000, + "burst_limit": 5, + "cooldown_period": 60, + "rate_limit_key": "user_id" + }, + "examples": [ + { + "description": "Basic descriptive analysis", + "input": { + "data": { + "columns": [ + "age", + "salary", + "department" + ], + "rows": [ + [ + 25, + 50000, + "engineering" + ], + [ + 30, + 60000, + "engineering" + ], + [ + 28, + 55000, + "marketing" + ] + ] + }, + "analysis_type": "descriptive", + "target_column": "salary" + }, + "expected_output": { + "insights": [ + "Average salary is $55,000", + "Salary range: $50,000 - $60,000", + "Engineering department has higher average salary" + ], + "statistics": { + "mean": 55000, + "median": 55000, + "std_dev": 5000 + } + } + } + ], + "metadata": { + "category": "data", + "idempotent": true, + "side_effects": [ + "May create temporary analysis files", + "Logs analysis parameters for optimization" + ], + "dependencies": [ + "statistics_engine", + "visualization_service" + ], + "security_requirements": [ + "Data anonymization", + "Access control validation" + ], + "generated_at": "2024-01-15T10:30:00Z", + "schema_version": "1.0", + "input_parameters": 4, + "output_parameters": 3, + "required_parameters": 2, + "optional_parameters": 2 + } + } + ], + "metadata": { + "generated_by": "tool_schema_generator.py", + "input_file": "sample_tool_descriptions.json", + "tool_count": 2, + "generation_timestamp": "2024-01-15T10:30:00Z", + "schema_version": "1.0" + }, + "validation_summary": { + "total_tools": 2, + "total_parameters": 8, + "total_validation_rules": 3, + "total_examples": 2 + } +} \ No newline at end of file diff --git a/skills/agent-designer/references/agent_architecture_patterns.md b/skills/agent-designer/references/agent_architecture_patterns.md new file mode 100644 index 00000000..cfa85ff5 --- /dev/null +++ b/skills/agent-designer/references/agent_architecture_patterns.md @@ -0,0 +1,445 @@ +# Agent Architecture Patterns Catalog + +## Overview + +This document provides a comprehensive catalog of multi-agent system architecture patterns, their characteristics, use cases, and implementation considerations. + +## Pattern Categories + +### 1. Single Agent Pattern + +**Description:** One agent handles all system functionality +**Structure:** User → Agent ← Tools +**Complexity:** Low + +**Characteristics:** +- Centralized decision making +- No inter-agent communication +- Simple state management +- Direct user interaction + +**Use Cases:** +- Personal assistants +- Simple automation tasks +- Prototyping and development +- Domain-specific applications + +**Advantages:** +- Simple to implement and debug +- Predictable behavior +- Low coordination overhead +- Clear responsibility model + +**Disadvantages:** +- Limited scalability +- Single point of failure +- Resource bottlenecks +- Difficulty handling complex workflows + +**Implementation Patterns:** +``` +Agent { + receive_request() + process_task() + use_tools() + return_response() +} +``` + +### 2. Supervisor Pattern (Hierarchical Delegation) + +**Description:** One supervisor coordinates multiple specialist agents +**Structure:** User → Supervisor → Specialists +**Complexity:** Medium + +**Characteristics:** +- Central coordination +- Clear hierarchy +- Specialized capabilities +- Delegation and aggregation + +**Use Cases:** +- Task decomposition scenarios +- Quality control workflows +- Resource allocation systems +- Project management + +**Advantages:** +- Clear command structure +- Specialized expertise +- Centralized quality control +- Efficient resource allocation + +**Disadvantages:** +- Supervisor bottleneck +- Complex coordination logic +- Single point of failure +- Limited parallelism + +**Implementation Patterns:** +``` +Supervisor { + decompose_task() + delegate_to_specialists() + monitor_progress() + aggregate_results() + quality_control() +} + +Specialist { + receive_assignment() + execute_specialized_task() + report_results() +} +``` + +### 3. Swarm Pattern (Peer-to-Peer) + +**Description:** Multiple autonomous agents collaborate as peers +**Structure:** Agent ↔ Agent ↔ Agent (interconnected) +**Complexity:** High + +**Characteristics:** +- Distributed decision making +- Peer-to-peer communication +- Emergent behavior +- Self-organization + +**Use Cases:** +- Distributed problem solving +- Parallel processing +- Fault-tolerant systems +- Research and exploration + +**Advantages:** +- High fault tolerance +- Scalable parallelism +- Emergent intelligence +- No single point of failure + +**Disadvantages:** +- Complex coordination +- Unpredictable behavior +- Difficult debugging +- Consensus overhead + +**Implementation Patterns:** +``` +SwarmAgent { + discover_peers() + share_information() + negotiate_tasks() + collaborate() + adapt_behavior() +} + +ConsensusProtocol { + propose_action() + vote() + reach_agreement() + execute_collective_decision() +} +``` + +### 4. Hierarchical Pattern (Multi-Level Management) + +**Description:** Multiple levels of management and execution +**Structure:** Executive → Managers → Workers (tree structure) +**Complexity:** Very High + +**Characteristics:** +- Multi-level hierarchy +- Distributed management +- Clear organizational structure +- Scalable command structure + +**Use Cases:** +- Enterprise systems +- Large-scale operations +- Complex workflows +- Organizational modeling + +**Advantages:** +- Natural organizational mapping +- Scalable structure +- Clear responsibilities +- Efficient resource management + +**Disadvantages:** +- Communication overhead +- Multi-level bottlenecks +- Complex coordination +- Slower decision making + +**Implementation Patterns:** +``` +Executive { + strategic_planning() + resource_allocation() + performance_monitoring() +} + +Manager { + tactical_planning() + team_coordination() + progress_reporting() +} + +Worker { + task_execution() + status_reporting() + resource_requests() +} +``` + +### 5. Pipeline Pattern (Sequential Processing) + +**Description:** Agents arranged in processing pipeline +**Structure:** Input → Stage1 → Stage2 → Stage3 → Output +**Complexity:** Medium + +**Characteristics:** +- Sequential processing +- Specialized stages +- Data flow architecture +- Clear processing order + +**Use Cases:** +- Data processing pipelines +- Manufacturing workflows +- Content processing +- ETL operations + +**Advantages:** +- Clear data flow +- Specialized optimization +- Predictable processing +- Easy to scale stages + +**Disadvantages:** +- Sequential bottlenecks +- Rigid processing order +- Stage coupling +- Limited flexibility + +**Implementation Patterns:** +``` +PipelineStage { + receive_input() + process_data() + validate_output() + send_to_next_stage() +} + +PipelineController { + manage_flow() + handle_errors() + monitor_throughput() + optimize_stages() +} +``` + +## Pattern Selection Criteria + +### Team Size Considerations +- **1 Agent:** Single Agent Pattern only +- **2-5 Agents:** Supervisor, Pipeline +- **6-15 Agents:** Swarm, Hierarchical, Pipeline +- **15+ Agents:** Hierarchical, Large Swarm + +### Task Complexity +- **Simple:** Single Agent +- **Medium:** Supervisor, Pipeline +- **Complex:** Swarm, Hierarchical +- **Very Complex:** Hierarchical + +### Coordination Requirements +- **None:** Single Agent +- **Low:** Pipeline, Supervisor +- **Medium:** Hierarchical +- **High:** Swarm + +### Fault Tolerance Requirements +- **Low:** Single Agent, Pipeline +- **Medium:** Supervisor, Hierarchical +- **High:** Swarm + +## Hybrid Patterns + +### Hub-and-Spoke with Clusters +Combines supervisor pattern with swarm clusters +- Central coordinator +- Specialized swarm clusters +- Hierarchical communication + +### Pipeline with Parallel Stages +Pipeline stages that can process in parallel +- Sequential overall flow +- Parallel processing within stages +- Load balancing across stage instances + +### Hierarchical Swarms +Swarm behavior at each hierarchical level +- Distributed decision making +- Hierarchical coordination +- Multi-level autonomy + +## Communication Patterns by Architecture + +### Single Agent +- Direct user interface +- Tool API calls +- No inter-agent communication + +### Supervisor +- Command/response with specialists +- Progress reporting +- Result aggregation + +### Swarm +- Broadcast messaging +- Peer discovery +- Consensus protocols +- Information sharing + +### Hierarchical +- Upward reporting +- Downward delegation +- Lateral coordination +- Skip-level communication + +### Pipeline +- Stage-to-stage data flow +- Error propagation +- Status monitoring +- Flow control + +## Scaling Considerations + +### Horizontal Scaling +- **Single Agent:** Scale by replication +- **Supervisor:** Scale specialists +- **Swarm:** Add more peers +- **Hierarchical:** Add at appropriate levels +- **Pipeline:** Scale bottleneck stages + +### Vertical Scaling +- **Single Agent:** More powerful agent +- **Supervisor:** Enhanced supervisor capabilities +- **Swarm:** Smarter individual agents +- **Hierarchical:** Better management agents +- **Pipeline:** Optimize stage processing + +## Error Handling Patterns + +### Single Agent +- Retry logic +- Fallback behaviors +- User notification + +### Supervisor +- Specialist failure detection +- Task reassignment +- Result validation + +### Swarm +- Peer failure detection +- Consensus recalculation +- Self-healing behavior + +### Hierarchical +- Escalation procedures +- Skip-level communication +- Management override + +### Pipeline +- Stage failure recovery +- Data replay +- Circuit breakers + +## Performance Characteristics + +| Pattern | Latency | Throughput | Scalability | Reliability | Complexity | +|---------|---------|------------|-------------|-------------|------------| +| Single Agent | Low | Low | Poor | Poor | Low | +| Supervisor | Medium | Medium | Good | Medium | Medium | +| Swarm | High | High | Excellent | Excellent | High | +| Hierarchical | Medium | High | Excellent | Good | Very High | +| Pipeline | Low | High | Good | Medium | Medium | + +## Best Practices by Pattern + +### Single Agent +- Keep scope focused +- Implement comprehensive error handling +- Use efficient tool selection +- Monitor resource usage + +### Supervisor +- Design clear delegation rules +- Implement progress monitoring +- Use timeout mechanisms +- Plan for specialist failures + +### Swarm +- Design simple interaction protocols +- Implement conflict resolution +- Monitor emergent behavior +- Plan for network partitions + +### Hierarchical +- Define clear role boundaries +- Implement efficient communication +- Plan escalation procedures +- Monitor span of control + +### Pipeline +- Optimize bottleneck stages +- Implement error recovery +- Use appropriate buffering +- Monitor flow rates + +## Anti-Patterns to Avoid + +### God Agent +Single agent that tries to do everything +- Violates single responsibility +- Creates maintenance nightmare +- Poor scalability + +### Chatty Communication +Excessive inter-agent messaging +- Performance degradation +- Network congestion +- Poor scalability + +### Circular Dependencies +Agents depending on each other cyclically +- Deadlock potential +- Complex error handling +- Difficult debugging + +### Over-Centralization +Too much logic in coordinator +- Single point of failure +- Bottleneck creation +- Poor fault tolerance + +### Under-Specification +Unclear roles and responsibilities +- Coordination failures +- Duplicate work +- Inconsistent behavior + +## Conclusion + +The choice of agent architecture pattern depends on multiple factors including team size, task complexity, coordination requirements, fault tolerance needs, and performance objectives. Each pattern has distinct trade-offs that must be carefully considered in the context of specific system requirements. + +Success factors include: +- Clear role definitions +- Appropriate communication patterns +- Robust error handling +- Scalability planning +- Performance monitoring + +The patterns can be combined and customized to meet specific needs, but maintaining clarity and avoiding unnecessary complexity should always be prioritized. \ No newline at end of file diff --git a/skills/agent-designer/references/evaluation_methodology.md b/skills/agent-designer/references/evaluation_methodology.md new file mode 100644 index 00000000..3b430f5b --- /dev/null +++ b/skills/agent-designer/references/evaluation_methodology.md @@ -0,0 +1,749 @@ +# Multi-Agent System Evaluation Methodology + +## Overview + +This document provides a comprehensive methodology for evaluating multi-agent systems across multiple dimensions including performance, reliability, cost-effectiveness, and user satisfaction. The methodology is designed to provide actionable insights for system optimization. + +## Evaluation Framework + +### Evaluation Dimensions + +#### 1. Task Performance +- **Success Rate:** Percentage of tasks completed successfully +- **Completion Time:** Time from task initiation to completion +- **Quality Metrics:** Accuracy, relevance, completeness of results +- **Partial Success:** Progress made on incomplete tasks + +#### 2. System Reliability +- **Availability:** System uptime and accessibility +- **Error Rates:** Frequency and types of errors +- **Recovery Time:** Time to recover from failures +- **Fault Tolerance:** System behavior under component failures + +#### 3. Cost Efficiency +- **Resource Utilization:** CPU, memory, network, storage usage +- **Token Consumption:** LLM API usage and costs +- **Operational Costs:** Infrastructure and maintenance costs +- **Cost per Task:** Economic efficiency per completed task + +#### 4. User Experience +- **Response Time:** User-perceived latency +- **User Satisfaction:** Qualitative feedback scores +- **Usability:** Ease of system interaction +- **Predictability:** Consistency of system behavior + +#### 5. Scalability +- **Load Handling:** Performance under increasing load +- **Resource Scaling:** Ability to scale resources dynamically +- **Concurrency:** Handling multiple simultaneous requests +- **Degradation Patterns:** Behavior at capacity limits + +#### 6. Security +- **Access Control:** Authentication and authorization effectiveness +- **Data Protection:** Privacy and confidentiality measures +- **Audit Trail:** Logging and monitoring completeness +- **Vulnerability Assessment:** Security weakness identification + +## Metrics Collection + +### Core Metrics + +#### Performance Metrics +```json +{ + "task_metrics": { + "task_id": "string", + "agent_id": "string", + "task_type": "string", + "start_time": "ISO 8601 timestamp", + "end_time": "ISO 8601 timestamp", + "duration_ms": "integer", + "status": "success|failure|partial|timeout", + "quality_score": "float 0-1", + "steps_completed": "integer", + "total_steps": "integer" + } +} +``` + +#### Resource Metrics +```json +{ + "resource_metrics": { + "timestamp": "ISO 8601 timestamp", + "agent_id": "string", + "cpu_usage_percent": "float", + "memory_usage_mb": "integer", + "network_bytes_sent": "integer", + "network_bytes_received": "integer", + "tokens_consumed": "integer", + "api_calls_made": "integer" + } +} +``` + +#### Error Metrics +```json +{ + "error_metrics": { + "timestamp": "ISO 8601 timestamp", + "error_type": "string", + "error_code": "string", + "error_message": "string", + "agent_id": "string", + "task_id": "string", + "severity": "critical|high|medium|low", + "recovery_action": "string", + "resolved": "boolean" + } +} +``` + +### Advanced Metrics + +#### Agent Collaboration Metrics +```json +{ + "collaboration_metrics": { + "timestamp": "ISO 8601 timestamp", + "initiating_agent": "string", + "target_agent": "string", + "interaction_type": "request|response|broadcast|delegate", + "latency_ms": "integer", + "success": "boolean", + "payload_size_bytes": "integer", + "context_shared": "boolean" + } +} +``` + +#### Tool Usage Metrics +```json +{ + "tool_metrics": { + "timestamp": "ISO 8601 timestamp", + "agent_id": "string", + "tool_name": "string", + "invocation_duration_ms": "integer", + "success": "boolean", + "error_type": "string|null", + "input_size_bytes": "integer", + "output_size_bytes": "integer", + "cached_result": "boolean" + } +} +``` + +## Evaluation Methods + +### 1. Synthetic Benchmarks + +#### Task Complexity Levels +- **Level 1 (Simple):** Single-agent, single-tool tasks +- **Level 2 (Moderate):** Multi-tool tasks requiring coordination +- **Level 3 (Complex):** Multi-agent collaborative tasks +- **Level 4 (Advanced):** Long-running, multi-stage workflows +- **Level 5 (Expert):** Adaptive tasks requiring learning + +#### Benchmark Task Categories +```yaml +benchmark_categories: + information_retrieval: + - simple_web_search + - multi_source_research + - fact_verification + - comparative_analysis + + content_generation: + - text_summarization + - creative_writing + - technical_documentation + - multilingual_translation + + data_processing: + - data_cleaning + - statistical_analysis + - visualization_creation + - report_generation + + problem_solving: + - algorithm_development + - optimization_tasks + - troubleshooting + - decision_support + + workflow_automation: + - multi_step_processes + - conditional_workflows + - exception_handling + - resource_coordination +``` + +#### Benchmark Execution +```python +def run_benchmark_suite(agents, benchmark_tasks): + results = {} + + for category, tasks in benchmark_tasks.items(): + category_results = [] + + for task in tasks: + task_result = execute_benchmark_task( + agents=agents, + task=task, + timeout=task.max_duration, + repetitions=task.repetitions + ) + category_results.append(task_result) + + results[category] = analyze_category_results(category_results) + + return generate_benchmark_report(results) +``` + +### 2. A/B Testing + +#### Test Design +```yaml +ab_test_design: + hypothesis: "New agent architecture improves task success rate" + success_metrics: + primary: "task_success_rate" + secondary: ["response_time", "cost_per_task", "user_satisfaction"] + + test_configuration: + control_group: "current_architecture" + treatment_group: "new_architecture" + traffic_split: 50/50 + duration_days: 14 + minimum_sample_size: 1000 + + statistical_parameters: + confidence_level: 0.95 + minimum_detectable_effect: 0.05 + statistical_power: 0.8 +``` + +#### Analysis Framework +```python +def analyze_ab_test(control_data, treatment_data, metrics): + results = {} + + for metric in metrics: + control_values = extract_metric_values(control_data, metric) + treatment_values = extract_metric_values(treatment_data, metric) + + # Statistical significance test + stat_result = perform_statistical_test( + control_values, + treatment_values, + test_type=determine_test_type(metric) + ) + + # Effect size calculation + effect_size = calculate_effect_size( + control_values, + treatment_values + ) + + results[metric] = { + "control_mean": np.mean(control_values), + "treatment_mean": np.mean(treatment_values), + "p_value": stat_result.p_value, + "confidence_interval": stat_result.confidence_interval, + "effect_size": effect_size, + "practical_significance": assess_practical_significance( + effect_size, metric + ) + } + + return results +``` + +### 3. Load Testing + +#### Load Test Scenarios +```yaml +load_test_scenarios: + baseline_load: + concurrent_users: 10 + ramp_up_time: "5 minutes" + duration: "30 minutes" + + normal_load: + concurrent_users: 100 + ramp_up_time: "10 minutes" + duration: "1 hour" + + peak_load: + concurrent_users: 500 + ramp_up_time: "15 minutes" + duration: "2 hours" + + stress_test: + concurrent_users: 1000 + ramp_up_time: "20 minutes" + duration: "1 hour" + + spike_test: + phases: + - users: 100, duration: "10 minutes" + - users: 1000, duration: "5 minutes" # Spike + - users: 100, duration: "15 minutes" +``` + +#### Performance Thresholds +```yaml +performance_thresholds: + response_time: + p50: 2000ms # 50th percentile + p90: 5000ms # 90th percentile + p95: 8000ms # 95th percentile + p99: 15000ms # 99th percentile + + throughput: + minimum: 10 # requests per second + target: 50 # requests per second + + error_rate: + maximum: 5% # percentage of failed requests + + resource_utilization: + cpu_max: 80% + memory_max: 85% + network_max: 70% +``` + +### 4. Real-World Evaluation + +#### Production Monitoring +```yaml +production_metrics: + business_metrics: + - task_completion_rate + - user_retention_rate + - feature_adoption_rate + - time_to_value + + technical_metrics: + - system_availability + - mean_time_to_recovery + - resource_efficiency + - cost_per_transaction + + user_experience_metrics: + - net_promoter_score + - user_satisfaction_rating + - task_abandonment_rate + - help_desk_ticket_volume +``` + +#### Continuous Evaluation Pipeline +```python +class ContinuousEvaluationPipeline: + def __init__(self, metrics_collector, analyzer, alerting): + self.metrics_collector = metrics_collector + self.analyzer = analyzer + self.alerting = alerting + + def run_evaluation_cycle(self): + # Collect recent metrics + metrics = self.metrics_collector.collect_recent_metrics( + time_window="1 hour" + ) + + # Analyze performance + analysis = self.analyzer.analyze_metrics(metrics) + + # Check for anomalies + anomalies = self.analyzer.detect_anomalies( + metrics, + baseline_window="24 hours" + ) + + # Generate alerts if needed + if anomalies: + self.alerting.send_alerts(anomalies) + + # Update performance baselines + self.analyzer.update_baselines(metrics) + + return analysis +``` + +## Analysis Techniques + +### 1. Statistical Analysis + +#### Descriptive Statistics +```python +def calculate_descriptive_stats(data): + return { + "count": len(data), + "mean": np.mean(data), + "median": np.median(data), + "std_dev": np.std(data), + "min": np.min(data), + "max": np.max(data), + "percentiles": { + "p25": np.percentile(data, 25), + "p50": np.percentile(data, 50), + "p75": np.percentile(data, 75), + "p90": np.percentile(data, 90), + "p95": np.percentile(data, 95), + "p99": np.percentile(data, 99) + } + } +``` + +#### Correlation Analysis +```python +def analyze_metric_correlations(metrics_df): + correlation_matrix = metrics_df.corr() + + # Identify strong correlations + strong_correlations = [] + for i in range(len(correlation_matrix.columns)): + for j in range(i + 1, len(correlation_matrix.columns)): + corr_value = correlation_matrix.iloc[i, j] + if abs(corr_value) > 0.7: # Strong correlation threshold + strong_correlations.append({ + "metric1": correlation_matrix.columns[i], + "metric2": correlation_matrix.columns[j], + "correlation": corr_value, + "strength": "strong" if abs(corr_value) > 0.8 else "moderate" + }) + + return strong_correlations +``` + +### 2. Trend Analysis + +#### Time Series Analysis +```python +def analyze_performance_trends(time_series_data, metric): + # Decompose time series + decomposition = seasonal_decompose( + time_series_data[metric], + model='additive', + period=24 # Daily seasonality + ) + + # Trend detection + trend_slope = calculate_trend_slope(decomposition.trend) + + # Seasonality detection + seasonal_patterns = identify_seasonal_patterns(decomposition.seasonal) + + # Anomaly detection + anomalies = detect_anomalies_isolation_forest(time_series_data[metric]) + + return { + "trend_direction": "increasing" if trend_slope > 0 else "decreasing" if trend_slope < 0 else "stable", + "trend_strength": abs(trend_slope), + "seasonal_patterns": seasonal_patterns, + "anomalies": anomalies, + "forecast": generate_forecast(time_series_data[metric], periods=24) + } +``` + +### 3. Comparative Analysis + +#### Multi-System Comparison +```python +def compare_systems(system_metrics_dict): + comparison_results = {} + + metrics_to_compare = [ + "success_rate", "average_response_time", + "cost_per_task", "error_rate" + ] + + for metric in metrics_to_compare: + metric_values = { + system: metrics[metric] + for system, metrics in system_metrics_dict.items() + } + + # Rank systems by metric + ranked_systems = sorted( + metric_values.items(), + key=lambda x: x[1], + reverse=(metric in ["success_rate"]) # Higher is better for some metrics + ) + + # Calculate relative performance + best_value = ranked_systems[0][1] + relative_performance = { + system: value / best_value if best_value > 0 else 0 + for system, value in metric_values.items() + } + + comparison_results[metric] = { + "rankings": ranked_systems, + "relative_performance": relative_performance, + "best_system": ranked_systems[0][0] + } + + return comparison_results +``` + +## Quality Assurance + +### 1. Data Quality Validation + +#### Data Completeness Checks +```python +def validate_data_completeness(metrics_data): + completeness_report = {} + + required_fields = [ + "timestamp", "task_id", "agent_id", + "duration_ms", "status", "success" + ] + + for field in required_fields: + missing_count = metrics_data[field].isnull().sum() + total_count = len(metrics_data) + completeness_percentage = (total_count - missing_count) / total_count * 100 + + completeness_report[field] = { + "completeness_percentage": completeness_percentage, + "missing_count": missing_count, + "status": "pass" if completeness_percentage >= 95 else "fail" + } + + return completeness_report +``` + +#### Data Consistency Checks +```python +def validate_data_consistency(metrics_data): + consistency_issues = [] + + # Check timestamp ordering + if not metrics_data['timestamp'].is_monotonic_increasing: + consistency_issues.append("Timestamps are not in chronological order") + + # Check duration consistency + duration_negative = (metrics_data['duration_ms'] < 0).sum() + if duration_negative > 0: + consistency_issues.append(f"Found {duration_negative} negative durations") + + # Check status-success consistency + success_status_mismatch = ( + (metrics_data['status'] == 'success') != metrics_data['success'] + ).sum() + if success_status_mismatch > 0: + consistency_issues.append(f"Found {success_status_mismatch} status-success mismatches") + + return consistency_issues +``` + +### 2. Evaluation Reliability + +#### Reproducibility Framework +```python +class ReproducibleEvaluation: + def __init__(self, config): + self.config = config + self.random_seed = config.get('random_seed', 42) + + def setup_environment(self): + # Set random seeds + random.seed(self.random_seed) + np.random.seed(self.random_seed) + + # Configure logging + self.setup_evaluation_logging() + + # Snapshot system state + self.snapshot_system_state() + + def run_evaluation(self, test_suite): + self.setup_environment() + + # Execute evaluation with full logging + results = self.execute_test_suite(test_suite) + + # Verify reproducibility + self.verify_reproducibility(results) + + return results +``` + +## Reporting Framework + +### 1. Executive Summary Report + +#### Key Performance Indicators +```yaml +kpi_dashboard: + overall_health_score: 85/100 + + performance: + task_success_rate: 94.2% + average_response_time: 2.3s + p95_response_time: 8.1s + + reliability: + system_uptime: 99.8% + error_rate: 2.1% + mean_recovery_time: 45s + + cost_efficiency: + cost_per_task: $0.05 + token_utilization: 78% + resource_efficiency: 82% + + user_satisfaction: + net_promoter_score: 42 + task_completion_rate: 89% + user_retention_rate: 76% +``` + +#### Trend Indicators +```yaml +trend_analysis: + performance_trends: + success_rate: "↗ +2.3% vs last month" + response_time: "↘ -15% vs last month" + error_rate: "→ stable vs last month" + + cost_trends: + total_cost: "↗ +8% vs last month" + cost_per_task: "↘ -5% vs last month" + efficiency: "↗ +12% vs last month" +``` + +### 2. Technical Deep-Dive Report + +#### Performance Analysis +```markdown +## Performance Analysis + +### Task Success Patterns +- **Overall Success Rate**: 94.2% (target: 95%) +- **By Task Type**: + - Simple tasks: 98.1% success + - Complex tasks: 87.4% success + - Multi-agent tasks: 91.2% success + +### Response Time Distribution +- **Median**: 1.8 seconds +- **95th Percentile**: 8.1 seconds +- **Peak Hours Impact**: +35% slower during 9-11 AM + +### Error Analysis +- **Top Error Types**: + 1. Timeout errors (34% of failures) + 2. Rate limit exceeded (28% of failures) + 3. Invalid input (19% of failures) +``` + +#### Resource Utilization +```markdown +## Resource Utilization + +### Compute Resources +- **CPU Utilization**: 45% average, 78% peak +- **Memory Usage**: 6.2GB average, 12.1GB peak +- **Network I/O**: 125 MB/s average + +### API Usage +- **Token Consumption**: 2.4M tokens/day +- **Cost Breakdown**: + - GPT-4: 68% of token costs + - GPT-3.5: 28% of token costs + - Other models: 4% of token costs +``` + +### 3. Actionable Recommendations + +#### Performance Optimization +```yaml +recommendations: + high_priority: + - title: "Reduce timeout error rate" + impact: "Could improve success rate by 2.1%" + effort: "Medium" + timeline: "2 weeks" + + - title: "Optimize complex task handling" + impact: "Could improve complex task success by 5%" + effort: "High" + timeline: "4 weeks" + + medium_priority: + - title: "Implement intelligent caching" + impact: "Could reduce costs by 15%" + effort: "Medium" + timeline: "3 weeks" +``` + +## Continuous Improvement Process + +### 1. Evaluation Cadence + +#### Regular Evaluation Schedule +```yaml +evaluation_schedule: + real_time: + frequency: "continuous" + metrics: ["error_rate", "response_time", "system_health"] + + hourly: + frequency: "every hour" + metrics: ["throughput", "resource_utilization", "user_activity"] + + daily: + frequency: "daily at 2 AM UTC" + metrics: ["success_rates", "cost_analysis", "user_satisfaction"] + + weekly: + frequency: "every Sunday" + metrics: ["trend_analysis", "comparative_analysis", "capacity_planning"] + + monthly: + frequency: "first Monday of month" + metrics: ["comprehensive_evaluation", "benchmark_testing", "strategic_review"] +``` + +### 2. Performance Baseline Management + +#### Baseline Update Process +```python +def update_performance_baselines(current_metrics, historical_baselines): + updated_baselines = {} + + for metric, current_value in current_metrics.items(): + historical_values = historical_baselines.get(metric, []) + historical_values.append(current_value) + + # Keep rolling window of last 30 days + historical_values = historical_values[-30:] + + # Calculate new baseline + baseline = { + "mean": np.mean(historical_values), + "std": np.std(historical_values), + "p95": np.percentile(historical_values, 95), + "trend": calculate_trend(historical_values) + } + + updated_baselines[metric] = baseline + + return updated_baselines +``` + +## Conclusion + +Effective evaluation of multi-agent systems requires a comprehensive, multi-dimensional approach that combines quantitative metrics with qualitative assessments. The methodology should be: + +1. **Comprehensive**: Cover all aspects of system performance +2. **Continuous**: Provide ongoing monitoring and evaluation +3. **Actionable**: Generate specific, implementable recommendations +4. **Adaptable**: Evolve with system changes and requirements +5. **Reliable**: Produce consistent, reproducible results + +Regular evaluation using this methodology will ensure multi-agent systems continue to meet user needs while optimizing for cost, performance, and reliability. \ No newline at end of file diff --git a/skills/agent-designer/references/tool_design_best_practices.md b/skills/agent-designer/references/tool_design_best_practices.md new file mode 100644 index 00000000..d4584d2d --- /dev/null +++ b/skills/agent-designer/references/tool_design_best_practices.md @@ -0,0 +1,470 @@ +# Tool Design Best Practices for Multi-Agent Systems + +## Overview + +This document outlines comprehensive best practices for designing tools that work effectively within multi-agent systems. Tools are the primary interface between agents and external capabilities, making their design critical for system success. + +## Core Principles + +### 1. Single Responsibility Principle +Each tool should have a clear, focused purpose: +- **Do one thing well:** Avoid multi-purpose tools that try to solve many problems +- **Clear boundaries:** Well-defined input/output contracts +- **Predictable behavior:** Consistent results for similar inputs +- **Easy to understand:** Purpose should be obvious from name and description + +### 2. Idempotency +Tools should produce consistent results: +- **Safe operations:** Read operations should never modify state +- **Repeatable operations:** Same input should yield same output (when possible) +- **State handling:** Clear semantics for state-modifying operations +- **Error recovery:** Failed operations should be safely retryable + +### 3. Composability +Tools should work well together: +- **Standard interfaces:** Consistent input/output formats +- **Minimal assumptions:** Don't assume specific calling contexts +- **Chain-friendly:** Output of one tool can be input to another +- **Modular design:** Tools can be combined in different ways + +### 4. Robustness +Tools should handle edge cases gracefully: +- **Input validation:** Comprehensive validation of all inputs +- **Error handling:** Graceful degradation on failures +- **Resource management:** Proper cleanup and resource management +- **Timeout handling:** Operations should have reasonable timeouts + +## Input Schema Design + +### Schema Structure +```json +{ + "type": "object", + "properties": { + "parameter_name": { + "type": "string", + "description": "Clear, specific description", + "examples": ["example1", "example2"], + "minLength": 1, + "maxLength": 1000 + } + }, + "required": ["parameter_name"], + "additionalProperties": false +} +``` + +### Parameter Guidelines + +#### Required vs Optional Parameters +- **Required parameters:** Essential for tool function +- **Optional parameters:** Provide additional control or customization +- **Default values:** Sensible defaults for optional parameters +- **Parameter groups:** Related parameters should be grouped logically + +#### Parameter Types +- **Primitives:** string, number, boolean for simple values +- **Arrays:** For lists of similar items +- **Objects:** For complex structured data +- **Enums:** For fixed sets of valid values +- **Unions:** When multiple types are acceptable + +#### Validation Rules +- **String validation:** + - Length constraints (minLength, maxLength) + - Pattern matching for formats (email, URL, etc.) + - Character set restrictions + - Content filtering for security + +- **Numeric validation:** + - Range constraints (minimum, maximum) + - Multiple restrictions (multipleOf) + - Precision requirements + - Special value handling (NaN, infinity) + +- **Array validation:** + - Size constraints (minItems, maxItems) + - Item type validation + - Uniqueness requirements + - Ordering requirements + +- **Object validation:** + - Required property enforcement + - Additional property policies + - Nested validation rules + - Dependency validation + +### Input Examples + +#### Good Example: +```json +{ + "name": "search_web", + "description": "Search the web for information", + "parameters": { + "type": "object", + "properties": { + "query": { + "type": "string", + "description": "Search query string", + "minLength": 1, + "maxLength": 500, + "examples": ["latest AI developments", "weather forecast"] + }, + "limit": { + "type": "integer", + "description": "Maximum number of results to return", + "minimum": 1, + "maximum": 100, + "default": 10 + }, + "language": { + "type": "string", + "description": "Language code for search results", + "enum": ["en", "es", "fr", "de"], + "default": "en" + } + }, + "required": ["query"], + "additionalProperties": false + } +} +``` + +#### Bad Example: +```json +{ + "name": "do_stuff", + "description": "Does various operations", + "parameters": { + "type": "object", + "properties": { + "data": { + "type": "string", + "description": "Some data" + } + }, + "additionalProperties": true + } +} +``` + +## Output Schema Design + +### Response Structure +```json +{ + "success": true, + "data": { + // Actual response data + }, + "metadata": { + "timestamp": "2024-01-15T10:30:00Z", + "execution_time_ms": 234, + "version": "1.0" + }, + "warnings": [], + "pagination": { + "total": 100, + "page": 1, + "per_page": 10, + "has_next": true + } +} +``` + +### Data Consistency +- **Predictable structure:** Same structure regardless of success/failure +- **Type consistency:** Same data types across different calls +- **Null handling:** Clear semantics for missing/null values +- **Empty responses:** Consistent handling of empty result sets + +### Metadata Inclusion +- **Execution time:** Performance monitoring +- **Timestamps:** Audit trails and debugging +- **Version information:** Compatibility tracking +- **Request identifiers:** Correlation and debugging + +## Error Handling + +### Error Response Structure +```json +{ + "success": false, + "error": { + "code": "INVALID_INPUT", + "message": "The provided query is too short", + "details": { + "field": "query", + "provided_length": 0, + "minimum_length": 1 + }, + "retry_after": null, + "documentation_url": "https://docs.example.com/errors#INVALID_INPUT" + }, + "request_id": "req_12345" +} +``` + +### Error Categories + +#### Client Errors (4xx equivalent) +- **INVALID_INPUT:** Malformed or invalid parameters +- **MISSING_PARAMETER:** Required parameter not provided +- **VALIDATION_ERROR:** Parameter fails validation rules +- **AUTHENTICATION_ERROR:** Invalid or missing credentials +- **PERMISSION_ERROR:** Insufficient permissions +- **RATE_LIMIT_ERROR:** Too many requests + +#### Server Errors (5xx equivalent) +- **INTERNAL_ERROR:** Unexpected server error +- **SERVICE_UNAVAILABLE:** Downstream service unavailable +- **TIMEOUT_ERROR:** Operation timed out +- **RESOURCE_EXHAUSTED:** Out of resources (memory, disk, etc.) +- **DEPENDENCY_ERROR:** External dependency failed + +#### Tool-Specific Errors +- **DATA_NOT_FOUND:** Requested data doesn't exist +- **FORMAT_ERROR:** Data in unexpected format +- **PROCESSING_ERROR:** Error during data processing +- **CONFIGURATION_ERROR:** Tool misconfiguration + +### Error Recovery Strategies + +#### Retry Logic +```json +{ + "retry_policy": { + "max_attempts": 3, + "backoff_strategy": "exponential", + "base_delay_ms": 1000, + "max_delay_ms": 30000, + "retryable_errors": [ + "TIMEOUT_ERROR", + "SERVICE_UNAVAILABLE", + "RATE_LIMIT_ERROR" + ] + } +} +``` + +#### Fallback Behaviors +- **Graceful degradation:** Partial results when possible +- **Alternative approaches:** Different methods to achieve same goal +- **Cached responses:** Return stale data if fresh data unavailable +- **Default responses:** Safe default when specific response impossible + +## Security Considerations + +### Input Sanitization +- **SQL injection prevention:** Parameterized queries +- **XSS prevention:** HTML encoding of outputs +- **Command injection prevention:** Input validation and sandboxing +- **Path traversal prevention:** Path validation and restrictions + +### Authentication and Authorization +- **API key management:** Secure storage and rotation +- **Token validation:** JWT validation and expiration +- **Permission checking:** Role-based access control +- **Audit logging:** Security event logging + +### Data Protection +- **PII handling:** Detection and protection of personal data +- **Encryption:** Data encryption in transit and at rest +- **Data retention:** Compliance with retention policies +- **Access logging:** Who accessed what data when + +## Performance Optimization + +### Response Time +- **Caching strategies:** Result caching for repeated requests +- **Connection pooling:** Reuse connections to external services +- **Async processing:** Non-blocking operations where possible +- **Resource optimization:** Efficient resource utilization + +### Throughput +- **Batch operations:** Support for bulk operations +- **Parallel processing:** Concurrent execution where safe +- **Load balancing:** Distribute load across instances +- **Resource scaling:** Auto-scaling based on demand + +### Resource Management +- **Memory usage:** Efficient memory allocation and cleanup +- **CPU optimization:** Avoid unnecessary computations +- **Network efficiency:** Minimize network round trips +- **Storage optimization:** Efficient data structures and storage + +## Testing Strategies + +### Unit Testing +```python +def test_search_web_valid_input(): + result = search_web("test query", limit=5) + assert result["success"] is True + assert len(result["data"]["results"]) <= 5 + +def test_search_web_invalid_input(): + result = search_web("", limit=5) + assert result["success"] is False + assert result["error"]["code"] == "INVALID_INPUT" +``` + +### Integration Testing +- **End-to-end workflows:** Complete user scenarios +- **External service mocking:** Mock external dependencies +- **Error simulation:** Simulate various error conditions +- **Performance testing:** Load and stress testing + +### Contract Testing +- **Schema validation:** Validate against defined schemas +- **Backward compatibility:** Ensure changes don't break clients +- **API versioning:** Test multiple API versions +- **Consumer-driven contracts:** Test from consumer perspective + +## Documentation + +### Tool Documentation Template +```markdown +# Tool Name + +## Description +Brief description of what the tool does. + +## Parameters +### Required Parameters +- `parameter_name` (type): Description + +### Optional Parameters +- `optional_param` (type, default: value): Description + +## Response +Description of response format and data. + +## Examples +### Basic Usage +Input: +```json +{ + "parameter_name": "value" +} +``` + +Output: +```json +{ + "success": true, + "data": {...} +} +``` + +## Error Codes +- `ERROR_CODE`: Description of when this error occurs +``` + +### API Documentation +- **OpenAPI/Swagger specs:** Machine-readable API documentation +- **Interactive examples:** Runnable examples in documentation +- **Code samples:** Examples in multiple programming languages +- **Changelog:** Version history and breaking changes + +## Versioning Strategy + +### Semantic Versioning +- **Major version:** Breaking changes +- **Minor version:** New features, backward compatible +- **Patch version:** Bug fixes, no new features + +### API Evolution +- **Deprecation policy:** How to deprecate old features +- **Migration guides:** Help users upgrade to new versions +- **Backward compatibility:** Support for old versions +- **Feature flags:** Gradual rollout of new features + +## Monitoring and Observability + +### Metrics Collection +- **Usage metrics:** Call frequency, success rates +- **Performance metrics:** Response times, throughput +- **Error metrics:** Error rates by type +- **Resource metrics:** CPU, memory, network usage + +### Logging +```json +{ + "timestamp": "2024-01-15T10:30:00Z", + "tool_name": "search_web", + "request_id": "req_12345", + "agent_id": "agent_001", + "input_hash": "abc123", + "execution_time_ms": 234, + "success": true, + "error_code": null +} +``` + +### Alerting +- **Error rate thresholds:** Alert on high error rates +- **Performance degradation:** Alert on slow responses +- **Resource exhaustion:** Alert on resource limits +- **Service availability:** Alert on service downtime + +## Common Anti-Patterns + +### Tool Design Anti-Patterns +- **God tools:** Tools that try to do everything +- **Chatty tools:** Tools that require many calls for simple tasks +- **Stateful tools:** Tools that maintain state between calls +- **Inconsistent interfaces:** Tools with different conventions + +### Error Handling Anti-Patterns +- **Silent failures:** Failing without proper error reporting +- **Generic errors:** Non-descriptive error messages +- **Inconsistent error formats:** Different error structures +- **No retry guidance:** Not indicating if operation is retryable + +### Performance Anti-Patterns +- **Synchronous everything:** Not using async operations where appropriate +- **No caching:** Repeatedly fetching same data +- **Resource leaks:** Not properly cleaning up resources +- **Unbounded operations:** Operations that can run indefinitely + +## Best Practices Checklist + +### Design Phase +- [ ] Single, clear purpose +- [ ] Well-defined input/output contracts +- [ ] Comprehensive input validation +- [ ] Idempotent operations where possible +- [ ] Error handling strategy defined + +### Implementation Phase +- [ ] Robust error handling +- [ ] Input sanitization +- [ ] Resource management +- [ ] Timeout handling +- [ ] Logging implementation + +### Testing Phase +- [ ] Unit tests for all functionality +- [ ] Integration tests with dependencies +- [ ] Error condition testing +- [ ] Performance testing +- [ ] Security testing + +### Documentation Phase +- [ ] Complete API documentation +- [ ] Usage examples +- [ ] Error code documentation +- [ ] Performance characteristics +- [ ] Security considerations + +### Deployment Phase +- [ ] Monitoring setup +- [ ] Alerting configuration +- [ ] Performance baselines +- [ ] Security reviews +- [ ] Operational runbooks + +## Conclusion + +Well-designed tools are the foundation of effective multi-agent systems. They should be reliable, secure, performant, and easy to use. Following these best practices will result in tools that agents can effectively compose to solve complex problems while maintaining system reliability and security. \ No newline at end of file diff --git a/skills/agent-designer/tool_schema_generator.py b/skills/agent-designer/tool_schema_generator.py new file mode 100644 index 00000000..d5a49ee5 --- /dev/null +++ b/skills/agent-designer/tool_schema_generator.py @@ -0,0 +1,978 @@ +#!/usr/bin/env python3 +""" +Tool Schema Generator - Generate structured tool schemas for AI agents + +Given a description of desired tools (name, purpose, inputs, outputs), generates +structured tool schemas compatible with OpenAI function calling format and +Anthropic tool use format. Includes: input validation rules, error response +formats, example calls, rate limit suggestions. + +Input: tool descriptions JSON +Output: tool schemas (OpenAI + Anthropic format) + validation rules + example usage +""" + +import json +import argparse +import sys +import re +from typing import Dict, List, Any, Optional, Union, Tuple +from dataclasses import dataclass, asdict +from enum import Enum + + +class ParameterType(Enum): + """Parameter types for tool schemas""" + STRING = "string" + INTEGER = "integer" + NUMBER = "number" + BOOLEAN = "boolean" + ARRAY = "array" + OBJECT = "object" + NULL = "null" + + +class ValidationRule(Enum): + """Validation rule types""" + REQUIRED = "required" + MIN_LENGTH = "min_length" + MAX_LENGTH = "max_length" + PATTERN = "pattern" + ENUM = "enum" + MINIMUM = "minimum" + MAXIMUM = "maximum" + MIN_ITEMS = "min_items" + MAX_ITEMS = "max_items" + UNIQUE_ITEMS = "unique_items" + FORMAT = "format" + + +@dataclass +class ParameterSpec: + """Parameter specification for tool inputs/outputs""" + name: str + type: ParameterType + description: str + required: bool = False + default: Any = None + validation_rules: Dict[str, Any] = None + examples: List[Any] = None + deprecated: bool = False + + +@dataclass +class ErrorSpec: + """Error specification for tool responses""" + error_code: str + error_message: str + http_status: int + retry_after: Optional[int] = None + details: Dict[str, Any] = None + + +@dataclass +class RateLimitSpec: + """Rate limiting specification""" + requests_per_minute: int + requests_per_hour: int + requests_per_day: int + burst_limit: int + cooldown_period: int + rate_limit_key: str = "user_id" + + +@dataclass +class ToolDescription: + """Input tool description""" + name: str + purpose: str + category: str + inputs: List[Dict[str, Any]] + outputs: List[Dict[str, Any]] + error_conditions: List[str] + side_effects: List[str] + idempotent: bool + rate_limits: Dict[str, Any] + dependencies: List[str] + examples: List[Dict[str, Any]] + security_requirements: List[str] + + +@dataclass +class ToolSchema: + """Complete tool schema with validation and examples""" + name: str + description: str + openai_schema: Dict[str, Any] + anthropic_schema: Dict[str, Any] + validation_rules: List[Dict[str, Any]] + error_responses: List[ErrorSpec] + rate_limits: RateLimitSpec + examples: List[Dict[str, Any]] + metadata: Dict[str, Any] + + +class ToolSchemaGenerator: + """Generate structured tool schemas from descriptions""" + + def __init__(self): + self.common_patterns = self._define_common_patterns() + self.format_validators = self._define_format_validators() + self.security_templates = self._define_security_templates() + + def _define_common_patterns(self) -> Dict[str, str]: + """Define common regex patterns for validation""" + return { + "email": r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$", + "url": r"^https?:\/\/(www\.)?[-a-zA-Z0-9@:%._\+~#=]{1,256}\.[a-zA-Z0-9()]{1,6}\b([-a-zA-Z0-9()@:%_\+.~#?&//=]*)$", + "uuid": r"^[0-9a-f]{8}-[0-9a-f]{4}-[1-5][0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}$", + "phone": r"^\+?1?[0-9]{10,15}$", + "ip_address": r"^(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)$", + "date": r"^\d{4}-\d{2}-\d{2}$", + "datetime": r"^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}(?:\.\d{3})?Z?$", + "slug": r"^[a-z0-9]+(?:-[a-z0-9]+)*$", + "semantic_version": r"^(?P0|[1-9]\d*)\.(?P0|[1-9]\d*)\.(?P0|[1-9]\d*)(?:-(?P(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*)(?:\.(?:0|[1-9]\d*|\d*[a-zA-Z-][0-9a-zA-Z-]*))*))?(?:\+(?P[0-9a-zA-Z-]+(?:\.[0-9a-zA-Z-]+)*))?$" + } + + def _define_format_validators(self) -> Dict[str, Dict[str, Any]]: + """Define format validators for common data types""" + return { + "email": { + "type": "string", + "format": "email", + "pattern": self.common_patterns["email"], + "min_length": 5, + "max_length": 254 + }, + "url": { + "type": "string", + "format": "uri", + "pattern": self.common_patterns["url"], + "min_length": 7, + "max_length": 2048 + }, + "uuid": { + "type": "string", + "format": "uuid", + "pattern": self.common_patterns["uuid"], + "min_length": 36, + "max_length": 36 + }, + "date": { + "type": "string", + "format": "date", + "pattern": self.common_patterns["date"], + "min_length": 10, + "max_length": 10 + }, + "datetime": { + "type": "string", + "format": "date-time", + "pattern": self.common_patterns["datetime"], + "min_length": 19, + "max_length": 30 + }, + "password": { + "type": "string", + "min_length": 8, + "max_length": 128, + "pattern": r"^(?=.*[a-z])(?=.*[A-Z])(?=.*\d)(?=.*[@$!%*?&])[A-Za-z\d@$!%*?&]" + } + } + + def _define_security_templates(self) -> Dict[str, Dict[str, Any]]: + """Define security requirement templates""" + return { + "authentication_required": { + "requires_auth": True, + "auth_methods": ["bearer_token", "api_key"], + "scope_required": ["read", "write"] + }, + "rate_limited": { + "rate_limits": { + "requests_per_minute": 60, + "requests_per_hour": 1000, + "burst_limit": 10 + } + }, + "input_sanitization": { + "sanitize_html": True, + "validate_sql_injection": True, + "escape_special_chars": True + }, + "output_validation": { + "validate_response_schema": True, + "filter_sensitive_data": True, + "content_type_validation": True + } + } + + def parse_tool_description(self, description: ToolDescription) -> ParameterSpec: + """Parse tool description into structured parameters""" + input_params = [] + output_params = [] + + # Parse input parameters + for input_spec in description.inputs: + param = self._parse_parameter_spec(input_spec) + input_params.append(param) + + # Parse output parameters + for output_spec in description.outputs: + param = self._parse_parameter_spec(output_spec) + output_params.append(param) + + return input_params, output_params + + def _parse_parameter_spec(self, param_spec: Dict[str, Any]) -> ParameterSpec: + """Parse individual parameter specification""" + name = param_spec.get("name", "") + type_str = param_spec.get("type", "string") + description = param_spec.get("description", "") + required = param_spec.get("required", False) + default = param_spec.get("default") + examples = param_spec.get("examples", []) + + # Parse parameter type + param_type = self._parse_parameter_type(type_str) + + # Generate validation rules + validation_rules = self._generate_validation_rules(param_spec, param_type) + + return ParameterSpec( + name=name, + type=param_type, + description=description, + required=required, + default=default, + validation_rules=validation_rules, + examples=examples + ) + + def _parse_parameter_type(self, type_str: str) -> ParameterType: + """Parse parameter type from string""" + type_mapping = { + "str": ParameterType.STRING, + "string": ParameterType.STRING, + "text": ParameterType.STRING, + "int": ParameterType.INTEGER, + "integer": ParameterType.INTEGER, + "float": ParameterType.NUMBER, + "number": ParameterType.NUMBER, + "bool": ParameterType.BOOLEAN, + "boolean": ParameterType.BOOLEAN, + "list": ParameterType.ARRAY, + "array": ParameterType.ARRAY, + "dict": ParameterType.OBJECT, + "object": ParameterType.OBJECT, + "null": ParameterType.NULL, + "none": ParameterType.NULL + } + + return type_mapping.get(type_str.lower(), ParameterType.STRING) + + def _generate_validation_rules(self, param_spec: Dict[str, Any], param_type: ParameterType) -> Dict[str, Any]: + """Generate validation rules for a parameter""" + rules = {} + + # Type-specific validation + if param_type == ParameterType.STRING: + rules.update(self._generate_string_validation(param_spec)) + elif param_type == ParameterType.INTEGER: + rules.update(self._generate_integer_validation(param_spec)) + elif param_type == ParameterType.NUMBER: + rules.update(self._generate_number_validation(param_spec)) + elif param_type == ParameterType.ARRAY: + rules.update(self._generate_array_validation(param_spec)) + elif param_type == ParameterType.OBJECT: + rules.update(self._generate_object_validation(param_spec)) + + # Common validation rules + if param_spec.get("required", False): + rules["required"] = True + + if "enum" in param_spec: + rules["enum"] = param_spec["enum"] + + if "pattern" in param_spec: + rules["pattern"] = param_spec["pattern"] + elif self._detect_format(param_spec.get("name", ""), param_spec.get("description", "")): + format_name = self._detect_format(param_spec.get("name", ""), param_spec.get("description", "")) + if format_name in self.format_validators: + rules.update(self.format_validators[format_name]) + + return rules + + def _generate_string_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]: + """Generate string-specific validation rules""" + rules = {} + + if "min_length" in param_spec: + rules["minLength"] = param_spec["min_length"] + elif "min_len" in param_spec: + rules["minLength"] = param_spec["min_len"] + else: + # Infer from description + desc = param_spec.get("description", "").lower() + if "password" in desc: + rules["minLength"] = 8 + elif "email" in desc: + rules["minLength"] = 5 + elif "name" in desc: + rules["minLength"] = 1 + + if "max_length" in param_spec: + rules["maxLength"] = param_spec["max_length"] + elif "max_len" in param_spec: + rules["maxLength"] = param_spec["max_len"] + else: + # Reasonable defaults + desc = param_spec.get("description", "").lower() + if "password" in desc: + rules["maxLength"] = 128 + elif "email" in desc: + rules["maxLength"] = 254 + elif "description" in desc or "content" in desc: + rules["maxLength"] = 10000 + elif "name" in desc or "title" in desc: + rules["maxLength"] = 255 + else: + rules["maxLength"] = 1000 + + return rules + + def _generate_integer_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]: + """Generate integer-specific validation rules""" + rules = {} + + if "minimum" in param_spec: + rules["minimum"] = param_spec["minimum"] + elif "min" in param_spec: + rules["minimum"] = param_spec["min"] + else: + # Infer from context + name = param_spec.get("name", "").lower() + desc = param_spec.get("description", "").lower() + if any(word in name + desc for word in ["count", "quantity", "amount", "size", "limit"]): + rules["minimum"] = 0 + elif "page" in name + desc: + rules["minimum"] = 1 + elif "port" in name + desc: + rules["minimum"] = 1 + rules["maximum"] = 65535 + + if "maximum" in param_spec: + rules["maximum"] = param_spec["maximum"] + elif "max" in param_spec: + rules["maximum"] = param_spec["max"] + + return rules + + def _generate_number_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]: + """Generate number-specific validation rules""" + rules = {} + + if "minimum" in param_spec: + rules["minimum"] = param_spec["minimum"] + if "maximum" in param_spec: + rules["maximum"] = param_spec["maximum"] + if "exclusive_minimum" in param_spec: + rules["exclusiveMinimum"] = param_spec["exclusive_minimum"] + if "exclusive_maximum" in param_spec: + rules["exclusiveMaximum"] = param_spec["exclusive_maximum"] + if "multiple_of" in param_spec: + rules["multipleOf"] = param_spec["multiple_of"] + + return rules + + def _generate_array_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]: + """Generate array-specific validation rules""" + rules = {} + + if "min_items" in param_spec: + rules["minItems"] = param_spec["min_items"] + elif "min_length" in param_spec: + rules["minItems"] = param_spec["min_length"] + else: + rules["minItems"] = 0 + + if "max_items" in param_spec: + rules["maxItems"] = param_spec["max_items"] + elif "max_length" in param_spec: + rules["maxItems"] = param_spec["max_length"] + else: + rules["maxItems"] = 1000 # Reasonable default + + if param_spec.get("unique_items", False): + rules["uniqueItems"] = True + + if "item_type" in param_spec: + rules["items"] = {"type": param_spec["item_type"]} + + return rules + + def _generate_object_validation(self, param_spec: Dict[str, Any]) -> Dict[str, Any]: + """Generate object-specific validation rules""" + rules = {} + + if "properties" in param_spec: + rules["properties"] = param_spec["properties"] + + if "required_properties" in param_spec: + rules["required"] = param_spec["required_properties"] + + if "additional_properties" in param_spec: + rules["additionalProperties"] = param_spec["additional_properties"] + else: + rules["additionalProperties"] = False + + if "min_properties" in param_spec: + rules["minProperties"] = param_spec["min_properties"] + + if "max_properties" in param_spec: + rules["maxProperties"] = param_spec["max_properties"] + + return rules + + def _detect_format(self, name: str, description: str) -> Optional[str]: + """Detect parameter format from name and description""" + combined = (name + " " + description).lower() + + format_indicators = { + "email": ["email", "e-mail", "email_address"], + "url": ["url", "uri", "link", "website", "endpoint"], + "uuid": ["uuid", "guid", "identifier", "id"], + "date": ["date", "birthday", "created_date", "modified_date"], + "datetime": ["datetime", "timestamp", "created_at", "updated_at"], + "password": ["password", "secret", "token", "api_key"] + } + + for format_name, indicators in format_indicators.items(): + if any(indicator in combined for indicator in indicators): + return format_name + + return None + + def generate_openai_schema(self, description: ToolDescription, input_params: List[ParameterSpec]) -> Dict[str, Any]: + """Generate OpenAI function calling schema""" + properties = {} + required = [] + + for param in input_params: + prop_def = { + "type": param.type.value, + "description": param.description + } + + # Add validation rules + if param.validation_rules: + prop_def.update(param.validation_rules) + + # Add examples + if param.examples: + prop_def["examples"] = param.examples + + # Add default value + if param.default is not None: + prop_def["default"] = param.default + + properties[param.name] = prop_def + + if param.required: + required.append(param.name) + + schema = { + "name": description.name, + "description": description.purpose, + "parameters": { + "type": "object", + "properties": properties, + "required": required, + "additionalProperties": False + } + } + + return schema + + def generate_anthropic_schema(self, description: ToolDescription, input_params: List[ParameterSpec]) -> Dict[str, Any]: + """Generate Anthropic tool use schema""" + input_schema = { + "type": "object", + "properties": {}, + "required": [] + } + + for param in input_params: + prop_def = { + "type": param.type.value, + "description": param.description + } + + # Add validation rules (Anthropic uses subset of JSON Schema) + if param.validation_rules: + # Filter to supported validation rules + supported_rules = ["minLength", "maxLength", "minimum", "maximum", "pattern", "enum", "items"] + for rule, value in param.validation_rules.items(): + if rule in supported_rules: + prop_def[rule] = value + + input_schema["properties"][param.name] = prop_def + + if param.required: + input_schema["required"].append(param.name) + + schema = { + "name": description.name, + "description": description.purpose, + "input_schema": input_schema + } + + return schema + + def generate_error_responses(self, description: ToolDescription) -> List[ErrorSpec]: + """Generate error response specifications""" + error_specs = [] + + # Common errors + common_errors = [ + { + "error_code": "invalid_input", + "error_message": "Invalid input parameters provided", + "http_status": 400, + "details": {"validation_errors": []} + }, + { + "error_code": "authentication_required", + "error_message": "Authentication required to access this tool", + "http_status": 401 + }, + { + "error_code": "insufficient_permissions", + "error_message": "Insufficient permissions to perform this operation", + "http_status": 403 + }, + { + "error_code": "rate_limit_exceeded", + "error_message": "Rate limit exceeded. Please try again later", + "http_status": 429, + "retry_after": 60 + }, + { + "error_code": "internal_error", + "error_message": "Internal server error occurred", + "http_status": 500 + }, + { + "error_code": "service_unavailable", + "error_message": "Service temporarily unavailable", + "http_status": 503, + "retry_after": 300 + } + ] + + # Add common errors + for error in common_errors: + error_specs.append(ErrorSpec(**error)) + + # Add tool-specific errors based on error conditions + for condition in description.error_conditions: + if "not found" in condition.lower(): + error_specs.append(ErrorSpec( + error_code="resource_not_found", + error_message=f"Requested resource not found: {condition}", + http_status=404 + )) + elif "timeout" in condition.lower(): + error_specs.append(ErrorSpec( + error_code="operation_timeout", + error_message=f"Operation timed out: {condition}", + http_status=408, + retry_after=30 + )) + elif "quota" in condition.lower() or "limit" in condition.lower(): + error_specs.append(ErrorSpec( + error_code="quota_exceeded", + error_message=f"Quota or limit exceeded: {condition}", + http_status=429, + retry_after=3600 + )) + elif "dependency" in condition.lower(): + error_specs.append(ErrorSpec( + error_code="dependency_failure", + error_message=f"Dependency service failure: {condition}", + http_status=502 + )) + + return error_specs + + def generate_rate_limits(self, description: ToolDescription) -> RateLimitSpec: + """Generate rate limiting specification""" + rate_limits = description.rate_limits + + # Default rate limits based on tool category + defaults = { + "search": {"rpm": 60, "rph": 1000, "rpd": 10000, "burst": 10}, + "data": {"rpm": 30, "rph": 500, "rpd": 5000, "burst": 5}, + "api": {"rpm": 100, "rph": 2000, "rpd": 20000, "burst": 20}, + "file": {"rpm": 120, "rph": 3000, "rpd": 30000, "burst": 30}, + "compute": {"rpm": 10, "rph": 100, "rpd": 1000, "burst": 3}, + "communication": {"rpm": 30, "rph": 300, "rpd": 3000, "burst": 5} + } + + category_defaults = defaults.get(description.category.lower(), defaults["api"]) + + return RateLimitSpec( + requests_per_minute=rate_limits.get("requests_per_minute", category_defaults["rpm"]), + requests_per_hour=rate_limits.get("requests_per_hour", category_defaults["rph"]), + requests_per_day=rate_limits.get("requests_per_day", category_defaults["rpd"]), + burst_limit=rate_limits.get("burst_limit", category_defaults["burst"]), + cooldown_period=rate_limits.get("cooldown_period", 60), + rate_limit_key=rate_limits.get("rate_limit_key", "user_id") + ) + + def generate_examples(self, description: ToolDescription, input_params: List[ParameterSpec]) -> List[Dict[str, Any]]: + """Generate usage examples""" + examples = [] + + # Use provided examples if available + if description.examples: + for example in description.examples: + examples.append(example) + + # Generate synthetic examples + if len(examples) == 0: + synthetic_example = self._generate_synthetic_example(description, input_params) + if synthetic_example: + examples.append(synthetic_example) + + # Ensure we have multiple examples showing different scenarios + if len(examples) == 1 and len(input_params) > 1: + # Generate minimal example + minimal_example = self._generate_minimal_example(description, input_params) + if minimal_example and minimal_example != examples[0]: + examples.append(minimal_example) + + return examples + + def _generate_synthetic_example(self, description: ToolDescription, input_params: List[ParameterSpec]) -> Dict[str, Any]: + """Generate a synthetic example based on parameter specifications""" + example_input = {} + + for param in input_params: + if param.examples: + example_input[param.name] = param.examples[0] + elif param.default is not None: + example_input[param.name] = param.default + else: + example_input[param.name] = self._generate_example_value(param) + + # Generate expected output based on tool purpose + expected_output = self._generate_example_output(description) + + return { + "description": f"Example usage of {description.name}", + "input": example_input, + "expected_output": expected_output + } + + def _generate_minimal_example(self, description: ToolDescription, input_params: List[ParameterSpec]) -> Dict[str, Any]: + """Generate minimal example with only required parameters""" + example_input = {} + + for param in input_params: + if param.required: + if param.examples: + example_input[param.name] = param.examples[0] + else: + example_input[param.name] = self._generate_example_value(param) + + if not example_input: + return None + + expected_output = self._generate_example_output(description) + + return { + "description": f"Minimal example of {description.name} with required parameters only", + "input": example_input, + "expected_output": expected_output + } + + def _generate_example_value(self, param: ParameterSpec) -> Any: + """Generate example value for a parameter""" + if param.type == ParameterType.STRING: + format_examples = { + "email": "user@example.com", + "url": "https://example.com", + "uuid": "123e4567-e89b-12d3-a456-426614174000", + "date": "2024-01-15", + "datetime": "2024-01-15T10:30:00Z" + } + + # Check for format in validation rules + if param.validation_rules and "format" in param.validation_rules: + format_type = param.validation_rules["format"] + if format_type in format_examples: + return format_examples[format_type] + + # Check for patterns or enum + if param.validation_rules: + if "enum" in param.validation_rules: + return param.validation_rules["enum"][0] + + # Generate based on name/description + name_lower = param.name.lower() + if "name" in name_lower: + return "example_name" + elif "query" in name_lower or "search" in name_lower: + return "search query" + elif "path" in name_lower: + return "/path/to/resource" + elif "message" in name_lower: + return "Example message" + else: + return "example_value" + + elif param.type == ParameterType.INTEGER: + if param.validation_rules: + min_val = param.validation_rules.get("minimum", 0) + max_val = param.validation_rules.get("maximum", 100) + return min(max(42, min_val), max_val) + return 42 + + elif param.type == ParameterType.NUMBER: + if param.validation_rules: + min_val = param.validation_rules.get("minimum", 0.0) + max_val = param.validation_rules.get("maximum", 100.0) + return min(max(42.5, min_val), max_val) + return 42.5 + + elif param.type == ParameterType.BOOLEAN: + return True + + elif param.type == ParameterType.ARRAY: + return ["item1", "item2"] + + elif param.type == ParameterType.OBJECT: + return {"key": "value"} + + else: + return None + + def _generate_example_output(self, description: ToolDescription) -> Dict[str, Any]: + """Generate example output based on tool description""" + category = description.category.lower() + + if category == "search": + return { + "results": [ + {"title": "Example Result 1", "url": "https://example.com/1", "snippet": "Example snippet..."}, + {"title": "Example Result 2", "url": "https://example.com/2", "snippet": "Another snippet..."} + ], + "total_count": 2 + } + elif category == "data": + return { + "data": [{"id": 1, "value": "example"}, {"id": 2, "value": "another"}], + "metadata": {"count": 2, "processed_at": "2024-01-15T10:30:00Z"} + } + elif category == "file": + return { + "success": True, + "file_path": "/path/to/file.txt", + "size": 1024, + "modified_at": "2024-01-15T10:30:00Z" + } + elif category == "api": + return { + "status": "success", + "data": {"result": "operation completed successfully"}, + "timestamp": "2024-01-15T10:30:00Z" + } + else: + return { + "success": True, + "message": f"{description.name} executed successfully", + "result": "example result" + } + + def generate_tool_schema(self, description: ToolDescription) -> ToolSchema: + """Generate complete tool schema""" + # Parse parameters + input_params, output_params = self.parse_tool_description(description) + + # Generate schemas + openai_schema = self.generate_openai_schema(description, input_params) + anthropic_schema = self.generate_anthropic_schema(description, input_params) + + # Generate validation rules + validation_rules = [] + for param in input_params: + if param.validation_rules: + validation_rules.append({ + "parameter": param.name, + "rules": param.validation_rules + }) + + # Generate error responses + error_responses = self.generate_error_responses(description) + + # Generate rate limits + rate_limits = self.generate_rate_limits(description) + + # Generate examples + examples = self.generate_examples(description, input_params) + + # Generate metadata + metadata = { + "category": description.category, + "idempotent": description.idempotent, + "side_effects": description.side_effects, + "dependencies": description.dependencies, + "security_requirements": description.security_requirements, + "generated_at": "2024-01-15T10:30:00Z", + "schema_version": "1.0", + "input_parameters": len(input_params), + "output_parameters": len(output_params), + "required_parameters": sum(1 for p in input_params if p.required), + "optional_parameters": sum(1 for p in input_params if not p.required) + } + + return ToolSchema( + name=description.name, + description=description.purpose, + openai_schema=openai_schema, + anthropic_schema=anthropic_schema, + validation_rules=validation_rules, + error_responses=error_responses, + rate_limits=rate_limits, + examples=examples, + metadata=metadata + ) + + +def main(): + parser = argparse.ArgumentParser(description="Tool Schema Generator for AI Agents") + parser.add_argument("input_file", help="JSON file with tool descriptions") + parser.add_argument("-o", "--output", help="Output file prefix (default: tool_schemas)") + parser.add_argument("--format", choices=["json", "both"], default="both", + help="Output format") + parser.add_argument("--validate", action="store_true", + help="Validate generated schemas") + + args = parser.parse_args() + + try: + # Load tool descriptions + with open(args.input_file, 'r') as f: + tools_data = json.load(f) + + # Parse tool descriptions + tool_descriptions = [] + for tool_data in tools_data.get("tools", []): + tool_desc = ToolDescription(**tool_data) + tool_descriptions.append(tool_desc) + + # Generate schemas + generator = ToolSchemaGenerator() + schemas = [] + + for description in tool_descriptions: + schema = generator.generate_tool_schema(description) + schemas.append(schema) + print(f"Generated schema for: {schema.name}") + + # Prepare output + output_data = { + "tool_schemas": [asdict(schema) for schema in schemas], + "metadata": { + "generated_by": "tool_schema_generator.py", + "input_file": args.input_file, + "tool_count": len(schemas), + "generation_timestamp": "2024-01-15T10:30:00Z", + "schema_version": "1.0" + }, + "validation_summary": { + "total_tools": len(schemas), + "total_parameters": sum(schema.metadata["input_parameters"] for schema in schemas), + "total_validation_rules": sum(len(schema.validation_rules) for schema in schemas), + "total_examples": sum(len(schema.examples) for schema in schemas) + } + } + + # Output files + output_prefix = args.output or "tool_schemas" + + if args.format in ["json", "both"]: + with open(f"{output_prefix}.json", 'w') as f: + json.dump(output_data, f, indent=2, default=str) + print(f"JSON output written to {output_prefix}.json") + + if args.format == "both": + # Generate separate files for different formats + + # OpenAI format + openai_schemas = { + "functions": [schema.openai_schema for schema in schemas] + } + with open(f"{output_prefix}_openai.json", 'w') as f: + json.dump(openai_schemas, f, indent=2) + print(f"OpenAI schemas written to {output_prefix}_openai.json") + + # Anthropic format + anthropic_schemas = { + "tools": [schema.anthropic_schema for schema in schemas] + } + with open(f"{output_prefix}_anthropic.json", 'w') as f: + json.dump(anthropic_schemas, f, indent=2) + print(f"Anthropic schemas written to {output_prefix}_anthropic.json") + + # Validation rules + validation_data = { + "validation_rules": {schema.name: schema.validation_rules for schema in schemas} + } + with open(f"{output_prefix}_validation.json", 'w') as f: + json.dump(validation_data, f, indent=2) + print(f"Validation rules written to {output_prefix}_validation.json") + + # Usage examples + examples_data = { + "examples": {schema.name: schema.examples for schema in schemas} + } + with open(f"{output_prefix}_examples.json", 'w') as f: + json.dump(examples_data, f, indent=2) + print(f"Usage examples written to {output_prefix}_examples.json") + + # Print summary + print(f"\nSchema Generation Summary:") + print(f"Tools processed: {len(schemas)}") + print(f"Total input parameters: {sum(schema.metadata['input_parameters'] for schema in schemas)}") + print(f"Total validation rules: {sum(len(schema.validation_rules) for schema in schemas)}") + print(f"Total examples generated: {sum(len(schema.examples) for schema in schemas)}") + + # Validation if requested + if args.validate: + print("\nValidation Results:") + for schema in schemas: + validation_errors = [] + + # Basic validation checks + if not schema.openai_schema.get("parameters", {}).get("properties"): + validation_errors.append("Missing input parameters") + + if not schema.examples: + validation_errors.append("No usage examples") + + if not schema.validation_rules: + validation_errors.append("No validation rules defined") + + if validation_errors: + print(f" {schema.name}: {', '.join(validation_errors)}") + else: + print(f" {schema.name}: ✓ Valid") + + except Exception as e: + print(f"Error: {e}", file=sys.stderr) + sys.exit(1) + + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/skills/agent-workflow-designer/SKILL.md b/skills/agent-workflow-designer/SKILL.md new file mode 100644 index 00000000..f1380c52 --- /dev/null +++ b/skills/agent-workflow-designer/SKILL.md @@ -0,0 +1,443 @@ +--- +name: "agent-workflow-designer" +description: "Agent Workflow Designer" +--- + +# Agent Workflow Designer + +**Tier:** POWERFUL +**Category:** Engineering +**Domain:** Multi-Agent Systems / AI Orchestration + +--- + +## Overview + +Design production-grade multi-agent orchestration systems. Covers five core patterns (sequential pipeline, parallel fan-out/fan-in, hierarchical delegation, event-driven, consensus), platform-specific implementations, handoff protocols, state management, error recovery, context window budgeting, and cost optimization. + +--- + +## Core Capabilities + +- Pattern selection guide for any orchestration requirement +- Handoff protocol templates (structured context passing) +- State management patterns for multi-agent workflows +- Error recovery and retry strategies +- Context window budget management +- Cost optimization strategies per platform +- Platform-specific configs: Claude Code Agent Teams, OpenClaw, CrewAI, AutoGen + +--- + +## When to Use + +- Building a multi-step AI pipeline that exceeds one agent's context capacity +- Parallelizing research, generation, or analysis tasks for speed +- Creating specialist agents with defined roles and handoff contracts +- Designing fault-tolerant AI workflows for production + +--- + +## Pattern Selection Guide + +``` +Is the task sequential (each step needs previous output)? + YES → Sequential Pipeline + NO → Can tasks run in parallel? + YES → Parallel Fan-out/Fan-in + NO → Is there a hierarchy of decisions? + YES → Hierarchical Delegation + NO → Is it event-triggered? + YES → Event-Driven + NO → Need consensus/validation? + YES → Consensus Pattern +``` + +--- + +## Pattern 1: Sequential Pipeline + +**Use when:** Each step depends on the previous output. Research → Draft → Review → Polish. + +```python +# sequential_pipeline.py +from dataclasses import dataclass +from typing import Callable, Any +import anthropic + +@dataclass +class PipelineStage: + name: "str" + system_prompt: str + input_key: str # what to take from state + output_key: str # what to write to state + model: str = "claude-3-5-sonnet-20241022" + max_tokens: int = 2048 + +class SequentialPipeline: + def __init__(self, stages: list[PipelineStage]): + self.stages = stages + self.client = anthropic.Anthropic() + + def run(self, initial_input: str) -> dict: + state = {"input": initial_input} + + for stage in self.stages: + print(f"[{stage.name}] Processing...") + + stage_input = state.get(stage.input_key, "") + + response = self.client.messages.create( + model=stage.model, + max_tokens=stage.max_tokens, + system=stage.system_prompt, + messages=[{"role": "user", "content": stage_input}], + ) + + state[stage.output_key] = response.content[0].text + state[f"{stage.name}_tokens"] = response.usage.input_tokens + response.usage.output_tokens + + print(f"[{stage.name}] Done. Tokens: {state[f'{stage.name}_tokens']}") + + return state + +# Example: Blog post pipeline +pipeline = SequentialPipeline([ + PipelineStage( + name="researcher", + system_prompt="You are a research specialist. Given a topic, produce a structured research brief with: key facts, statistics, expert perspectives, and controversy points.", + input_key="input", + output_key="research", + ), + PipelineStage( + name="writer", + system_prompt="You are a senior content writer. Using the research provided, write a compelling 800-word blog post with a clear hook, 3 main sections, and a strong CTA.", + input_key="research", + output_key="draft", + ), + PipelineStage( + name="editor", + system_prompt="You are a copy editor. Review the draft for: clarity, flow, grammar, and SEO. Return the improved version only, no commentary.", + input_key="draft", + output_key="final", + ), +]) +``` + +--- + +## Pattern 2: Parallel Fan-out / Fan-in + +**Use when:** Independent tasks that can run concurrently. Research 5 competitors simultaneously. + +```python +# parallel_fanout.py +import asyncio +import anthropic +from typing import Any + +async def run_agent(client, task_name: "str-system-str-user-str-model-str"claude-3-5-sonnet-20241022") -> dict: + """Single async agent call""" + loop = asyncio.get_event_loop() + + def _call(): + return client.messages.create( + model=model, + max_tokens=2048, + system=system, + messages=[{"role": "user", "content": user}], + ) + + response = await loop.run_in_executor(None, _call) + return { + "task": task_name, + "output": response.content[0].text, + "tokens": response.usage.input_tokens + response.usage.output_tokens, + } + +async def parallel_research(competitors: list[str], research_type: str) -> dict: + """Fan-out: research all competitors in parallel. Fan-in: synthesize results.""" + client = anthropic.Anthropic() + + # FAN-OUT: spawn parallel agent calls + tasks = [ + run_agent( + client, + task_name=competitor, + system=f"You are a competitive intelligence analyst. Research {competitor} and provide: pricing, key features, target market, and known weaknesses.", + user=f"Analyze {competitor} for comparison with our product in the {research_type} market.", + ) + for competitor in competitors + ] + + results = await asyncio.gather(*tasks, return_exceptions=True) + + # Handle failures gracefully + successful = [r for r in results if not isinstance(r, Exception)] + failed = [r for r in results if isinstance(r, Exception)] + + if failed: + print(f"Warning: {len(failed)} research tasks failed: {failed}") + + # FAN-IN: synthesize + combined_research = "\n\n".join([ + f"## {r['task']}\n{r['output']}" for r in successful + ]) + + synthesis = await run_agent( + client, + task_name="synthesizer", + system="You are a strategic analyst. Synthesize competitor research into a concise comparison matrix and strategic recommendations.", + user=f"Synthesize these competitor analyses:\n\n{combined_research}", + model="claude-3-5-sonnet-20241022", + ) + + return { + "individual_analyses": successful, + "synthesis": synthesis["output"], + "total_tokens": sum(r["tokens"] for r in successful) + synthesis["tokens"], + } +``` + +--- + +## Pattern 3: Hierarchical Delegation + +**Use when:** Complex tasks with subtask discovery. Orchestrator breaks down work, delegates to specialists. + +```python +# hierarchical_delegation.py +import json +import anthropic + +ORCHESTRATOR_SYSTEM = """You are an orchestration agent. Your job is to: +1. Analyze the user's request +2. Break it into subtasks +3. Assign each to the appropriate specialist agent +4. Collect results and synthesize + +Available specialists: +- researcher: finds facts, data, and information +- writer: creates content and documents +- coder: writes and reviews code +- analyst: analyzes data and produces insights + +Respond with a JSON plan: +{ + "subtasks": [ + {"id": "1", "agent": "researcher", "task": "...", "depends_on": []}, + {"id": "2", "agent": "writer", "task": "...", "depends_on": ["1"]} + ] +}""" + +SPECIALIST_SYSTEMS = { + "researcher": "You are a research specialist. Find accurate, relevant information and cite sources when possible.", + "writer": "You are a professional writer. Create clear, engaging content in the requested format.", + "coder": "You are a senior software engineer. Write clean, well-commented code with error handling.", + "analyst": "You are a data analyst. Provide structured analysis with evidence-backed conclusions.", +} + +class HierarchicalOrchestrator: + def __init__(self): + self.client = anthropic.Anthropic() + + def run(self, user_request: str) -> str: + # 1. Orchestrator creates plan + plan_response = self.client.messages.create( + model="claude-3-5-sonnet-20241022", + max_tokens=1024, + system=ORCHESTRATOR_SYSTEM, + messages=[{"role": "user", "content": user_request}], + ) + + plan = json.loads(plan_response.content[0].text) + results = {} + + # 2. Execute subtasks respecting dependencies + for subtask in self._topological_sort(plan["subtasks"]): + context = self._build_context(subtask, results) + specialist = SPECIALIST_SYSTEMS[subtask["agent"]] + + result = self.client.messages.create( + model="claude-3-5-sonnet-20241022", + max_tokens=2048, + system=specialist, + messages=[{"role": "user", "content": f"{context}\n\nTask: {subtask['task']}"}], + ) + results[subtask["id"]] = result.content[0].text + + # 3. Final synthesis + all_results = "\n\n".join([f"### {k}\n{v}" for k, v in results.items()]) + synthesis = self.client.messages.create( + model="claude-3-5-sonnet-20241022", + max_tokens=2048, + system="Synthesize the specialist outputs into a coherent final response.", + messages=[{"role": "user", "content": f"Original request: {user_request}\n\nSpecialist outputs:\n{all_results}"}], + ) + return synthesis.content[0].text + + def _build_context(self, subtask: dict, results: dict) -> str: + if not subtask.get("depends_on"): + return "" + deps = [f"Output from task {dep}:\n{results[dep]}" for dep in subtask["depends_on"] if dep in results] + return "Previous results:\n" + "\n\n".join(deps) if deps else "" + + def _topological_sort(self, subtasks: list) -> list: + # Simple ordered execution respecting depends_on + ordered, remaining = [], list(subtasks) + completed = set() + while remaining: + for task in remaining: + if all(dep in completed for dep in task.get("depends_on", [])): + ordered.append(task) + completed.add(task["id"]) + remaining.remove(task) + break + return ordered +``` + +--- + +## Handoff Protocol Template + +```python +# Standard handoff context format — use between all agents +@dataclass +class AgentHandoff: + """Structured context passed between agents in a workflow.""" + task_id: str + workflow_id: str + step_number: int + total_steps: int + + # What was done + previous_agent: str + previous_output: str + artifacts: dict # {"filename": "content"} for any files produced + + # What to do next + current_agent: str + current_task: str + constraints: list[str] # hard rules for this step + + # Metadata + context_budget_remaining: int # tokens left for this agent + cost_so_far_usd: float + + def to_prompt(self) -> str: + return f""" +# Agent Handoff — Step {self.step_number}/{self.total_steps} + +## Your Task +{self.current_task} + +## Constraints +{chr(10).join(f'- {c}' for c in self.constraints)} + +## Context from Previous Step ({self.previous_agent}) +{self.previous_output[:2000]}{"... [truncated]" if len(self.previous_output) > 2000 else ""} + +## Context Budget +You have approximately {self.context_budget_remaining} tokens remaining. Be concise. +""" +``` + +--- + +## Error Recovery Patterns + +```python +import time +from functools import wraps + +def with_retry(max_attempts=3, backoff_seconds=2, fallback_model=None): + """Decorator for agent calls with exponential backoff and model fallback.""" + def decorator(fn): + @wraps(fn) + def wrapper(*args, **kwargs): + last_error = None + for attempt in range(max_attempts): + try: + return fn(*args, **kwargs) + except Exception as e: + last_error = e + if attempt < max_attempts - 1: + wait = backoff_seconds * (2 ** attempt) + print(f"Attempt {attempt+1} failed: {e}. Retrying in {wait}s...") + time.sleep(wait) + + # Fall back to cheaper/faster model on rate limit + if fallback_model and "rate_limit" in str(e).lower(): + kwargs["model"] = fallback_model + raise last_error + return wrapper + return decorator + +@with_retry(max_attempts=3, fallback_model="claude-3-haiku-20240307") +def call_agent(model, system, user): + ... +``` + +--- + +## Context Window Budgeting + +```python +# Budget context across a multi-step pipeline +# Rule: never let any step consume more than 60% of remaining budget + +CONTEXT_LIMITS = { + "claude-3-5-sonnet-20241022": 200_000, + "gpt-4o": 128_000, +} + +class ContextBudget: + def __init__(self, model: str, reserve_pct: float = 0.2): + total = CONTEXT_LIMITS.get(model, 128_000) + self.total = total + self.reserve = int(total * reserve_pct) # keep 20% as buffer + self.used = 0 + + @property + def remaining(self): + return self.total - self.reserve - self.used + + def allocate(self, step_name: "str-requested-int-int" + allocated = min(requested, int(self.remaining * 0.6)) # max 60% of remaining + print(f"[Budget] {step_name}: allocated {allocated:,} tokens (remaining: {self.remaining:,})") + return allocated + + def consume(self, tokens_used: int): + self.used += tokens_used + +def truncate_to_budget(text: str, token_budget: int, chars_per_token: float = 4.0) -> str: + """Rough truncation — use tiktoken for precision.""" + char_budget = int(token_budget * chars_per_token) + if len(text) <= char_budget: + return text + return text[:char_budget] + "\n\n[... truncated to fit context budget ...]" +``` + +--- + +## Cost Optimization Strategies + +| Strategy | Savings | Tradeoff | +|---|---|---| +| Use Haiku for routing/classification | 85-90% | Slightly less nuanced judgment | +| Cache repeated system prompts | 50-90% | Requires prompt caching setup | +| Truncate intermediate outputs | 20-40% | May lose detail in handoffs | +| Batch similar tasks | 50% | Latency increases | +| Use Sonnet for most, Opus for final step only | 60-70% | Final quality may improve | +| Short-circuit on confidence threshold | 30-50% | Need confidence scoring | + +--- + +## Common Pitfalls + +- **Circular dependencies** — agents calling each other in loops; enforce DAG structure at design time +- **Context bleed** — passing entire previous output to every step; summarize or extract only what's needed +- **No timeout** — a stuck agent blocks the whole pipeline; always set max_tokens and wall-clock timeouts +- **Silent failures** — agent returns plausible but wrong output; add validation steps for critical paths +- **Ignoring cost** — 10 parallel Opus calls is $0.50 per workflow; model selection is a cost decision +- **Over-orchestration** — if a single prompt can do it, it should; only add agents when genuinely needed diff --git a/skills/agent-workflow-designer/_meta.json b/skills/agent-workflow-designer/_meta.json new file mode 100644 index 00000000..72ae123d --- /dev/null +++ b/skills/agent-workflow-designer/_meta.json @@ -0,0 +1,11 @@ +{ + "owner": "alirezarezvani", + "slug": "agent-workflow-designer", + "displayName": "agent-workflow-designer", + "latest": { + "version": "1.0.0", + "publishedAt": 1773242231297, + "commit": "https://github.com/openclaw/skills/commit/32b393a243c11966e4db38bf8b2a26e5e4add6a6" + }, + "history": [] +} diff --git a/skills/agile-product-owner/SKILL.md b/skills/agile-product-owner/SKILL.md new file mode 100644 index 00000000..410424a2 --- /dev/null +++ b/skills/agile-product-owner/SKILL.md @@ -0,0 +1,391 @@ +--- +name: "agile-product-owner" +description: Agile product ownership for backlog management and sprint execution. Covers user story writing, acceptance criteria, sprint planning, and velocity tracking. Use for writing user stories, creating acceptance criteria, planning sprints, estimating story points, breaking down epics, or prioritizing backlog. +triggers: + - write user story + - create acceptance criteria + - plan sprint + - estimate story points + - break down epic + - prioritize backlog + - sprint planning + - INVEST criteria + - Given When Then + - user story template + - sprint capacity + - velocity tracking +--- + +# Agile Product Owner + +Backlog management and sprint execution toolkit for product owners, including user story generation, acceptance criteria patterns, sprint planning, and velocity tracking. + +--- + +## Table of Contents + +- [User Story Generation Workflow](#user-story-generation-workflow) +- [Acceptance Criteria Patterns](#acceptance-criteria-patterns) +- [Epic Breakdown Workflow](#epic-breakdown-workflow) +- [Sprint Planning Workflow](#sprint-planning-workflow) +- [Backlog Prioritization](#backlog-prioritization) +- [Reference Documentation](#reference-documentation) +- [Tools](#tools) + +--- + +## User Story Generation Workflow + +Create INVEST-compliant user stories from requirements: + +1. Identify the persona (who benefits from this feature) +2. Define the action or capability needed +3. Articulate the benefit or value delivered +4. Write acceptance criteria using Given-When-Then +5. Estimate story points using Fibonacci scale +6. Validate against INVEST criteria +7. Add to backlog with priority +8. **Validation:** Story passes all INVEST criteria; acceptance criteria are testable + +### User Story Template + +``` +As a [persona], +I want to [action/capability], +So that [benefit/value]. +``` + +**Example:** +``` +As a marketing manager, +I want to export campaign reports to PDF, +So that I can share results with stakeholders who don't have system access. +``` + +### Story Types + +| Type | Template | Example | +|------|----------|---------| +| Feature | As a [persona], I want to [action] so that [benefit] | As a user, I want to filter search results so that I find items faster | +| Improvement | As a [persona], I need [capability] to [goal] | As a user, I need faster page loads to complete tasks without frustration | +| Bug Fix | As a [persona], I expect [behavior] when [condition] | As a user, I expect my cart to persist when I refresh the page | +| Enabler | As a developer, I need to [technical task] to enable [capability] | As a developer, I need to implement caching to enable instant search | + +### Persona Reference + +| Persona | Typical Needs | Context | +|---------|--------------|---------| +| End User | Efficiency, simplicity, reliability | Daily feature usage | +| Administrator | Control, visibility, security | System management | +| Power User | Automation, customization, shortcuts | Expert workflows | +| New User | Guidance, learning, safety | Onboarding | + +--- + +## Acceptance Criteria Patterns + +Write testable acceptance criteria using Given-When-Then format. + +### Given-When-Then Template + +``` +Given [precondition/context], +When [action/trigger], +Then [expected outcome]. +``` + +**Examples:** +``` +Given the user is logged in with valid credentials, +When they click the "Export" button, +Then a PDF download starts within 2 seconds. + +Given the user has entered an invalid email format, +When they submit the registration form, +Then an inline error message displays "Please enter a valid email address." + +Given the shopping cart contains items, +When the user refreshes the browser, +Then the cart contents remain unchanged. +``` + +### Acceptance Criteria Checklist + +Each story should include criteria for: + +| Category | Example | +|----------|---------| +| Happy Path | Given valid input, When submitted, Then success message displayed | +| Validation | Should reject input when required field is empty | +| Error Handling | Must show user-friendly message when API fails | +| Performance | Should complete operation within 2 seconds | +| Accessibility | Must be navigable via keyboard only | + +### Minimum Criteria by Story Size + +| Story Points | Minimum AC Count | +|--------------|------------------| +| 1-2 | 3-4 criteria | +| 3-5 | 4-6 criteria | +| 8 | 5-8 criteria | +| 13+ | Split the story | + +See `references/user-story-templates.md` for complete template library. + +--- + +## Epic Breakdown Workflow + +Break epics into deliverable sprint-sized stories: + +1. Define epic scope and success criteria +2. Identify all personas affected by the epic +3. List all capabilities needed for each persona +4. Group capabilities into logical stories +5. Validate each story is ≤8 points +6. Identify dependencies between stories +7. Sequence stories for incremental delivery +8. **Validation:** Each story delivers standalone value; total stories cover epic scope + +### Splitting Techniques + +| Technique | When to Use | Example | +|-----------|-------------|---------| +| By workflow step | Linear process | "Checkout" → "Add to cart" + "Enter payment" + "Confirm order" | +| By persona | Multiple user types | "Dashboard" → "Admin dashboard" + "User dashboard" | +| By data type | Multiple inputs | "Import" → "Import CSV" + "Import Excel" | +| By operation | CRUD functionality | "Manage users" → "Create" + "Edit" + "Delete" | +| Happy path first | Risk reduction | "Feature" → "Basic flow" + "Error handling" + "Edge cases" | + +### Epic Example + +**Epic:** User Dashboard + +**Breakdown:** +``` +Epic: User Dashboard (34 points total) +├── US-001: View key metrics (5 pts) - End User +├── US-002: Customize layout (5 pts) - Power User +├── US-003: Export data to CSV (3 pts) - End User +├── US-004: Share with team (5 pts) - End User +├── US-005: Set up alerts (5 pts) - Power User +├── US-006: Filter by date range (3 pts) - End User +├── US-007: Admin overview (5 pts) - Admin +└── US-008: Enable caching (3 pts) - Enabler +``` + +--- + +## Sprint Planning Workflow + +Plan sprint capacity and select stories: + +1. Calculate team capacity (velocity × availability) +2. Review sprint goal with stakeholders +3. Select stories from prioritized backlog +4. Fill to 80-85% of capacity (committed) +5. Add stretch goals (10-15% additional) +6. Identify dependencies and risks +7. Break complex stories into tasks +8. **Validation:** Committed points ≤85% capacity; all stories have acceptance criteria + +### Capacity Calculation + +``` +Sprint Capacity = Average Velocity × Availability Factor + +Example: +Average Velocity: 30 points +Team availability: 90% (one member partially out) +Adjusted Capacity: 27 points + +Committed: 23 points (85% of 27) +Stretch: 4 points (15% of 27) +``` + +### Availability Factors + +| Scenario | Factor | +|----------|--------| +| Full sprint, no PTO | 1.0 | +| One team member out 50% | 0.9 | +| Holiday during sprint | 0.8 | +| Multiple members out | 0.7 | + +### Sprint Loading Template + +``` +Sprint Capacity: 27 points +Sprint Goal: [Clear, measurable objective] + +COMMITTED (23 points): +[H] US-001: User dashboard (5 pts) +[H] US-002: Export feature (3 pts) +[H] US-003: Search filter (5 pts) +[M] US-004: Settings page (5 pts) +[M] US-005: Help tooltips (3 pts) +[L] US-006: Theme options (2 pts) + +STRETCH (4 points): +[L] US-007: Sort options (2 pts) +[L] US-008: Print view (2 pts) +``` + +See `references/sprint-planning-guide.md` for complete planning procedures. + +--- + +## Backlog Prioritization + +Prioritize backlog using value and effort assessment. + +### Priority Levels + +| Priority | Definition | Sprint Target | +|----------|------------|---------------| +| Critical | Blocking users, security, data loss | Immediate | +| High | Core functionality, key user needs | This sprint | +| Medium | Improvements, enhancements | Next 2-3 sprints | +| Low | Nice-to-have, minor improvements | Backlog | + +### Prioritization Factors + +| Factor | Weight | Questions | +|--------|--------|-----------| +| Business Value | 40% | Revenue impact? User demand? Strategic alignment? | +| User Impact | 30% | How many users? How frequently used? | +| Risk/Dependencies | 15% | Technical risk? External dependencies? | +| Effort | 15% | Size? Complexity? Uncertainty? | + +### INVEST Criteria Validation + +Before adding to sprint, validate each story: + +| Criterion | Question | Pass If... | +|-----------|----------|------------| +| **I**ndependent | Can this be developed without other uncommitted stories? | No blocking dependencies | +| **N**egotiable | Is the implementation flexible? | Multiple approaches possible | +| **V**aluable | Does this deliver user or business value? | Clear benefit in "so that" | +| **E**stimable | Can the team estimate this? | Understood well enough to size | +| **S**mall | Can this complete in one sprint? | ≤8 story points | +| **T**estable | Can we verify this is done? | Clear acceptance criteria | + +--- + +## Reference Documentation + +### User Story Templates + +`references/user-story-templates.md` contains: + +- Standard story formats by type (feature, improvement, bug fix, enabler) +- Acceptance criteria patterns (Given-When-Then, Should/Must/Can) +- INVEST criteria validation checklist +- Story point estimation guide (Fibonacci scale) +- Common story antipatterns and fixes +- Story splitting techniques + +### Sprint Planning Guide + +`references/sprint-planning-guide.md` contains: + +- Sprint planning meeting agenda +- Capacity calculation formulas +- Backlog prioritization framework (WSJF) +- Sprint ceremony guides (standup, review, retro) +- Velocity tracking and burndown patterns +- Definition of Done checklist +- Sprint metrics and targets + +--- + +## Tools + +### User Story Generator + +```bash +# Generate stories from sample epic +python scripts/user_story_generator.py + +# Plan sprint with capacity +python scripts/user_story_generator.py sprint 30 +``` + +Generates: +- INVEST-compliant user stories +- Given-When-Then acceptance criteria +- Story point estimates (Fibonacci scale) +- Priority assignments +- Sprint loading with committed and stretch items + +### Sample Output + +``` +USER STORY: USR-001 +======================================== +Title: View Key Metrics +Type: story +Priority: HIGH +Points: 5 + +Story: +As a End User, I want to view key metrics and KPIs +so that I can save time and work more efficiently + +Acceptance Criteria: + 1. Given user has access, When they view key metrics, Then the result is displayed + 2. Should validate input before processing + 3. Must show clear error message when action fails + 4. Should complete within 2 seconds + 5. Must be accessible via keyboard navigation + +INVEST Checklist: + ✓ Independent + ✓ Negotiable + ✓ Valuable + ✓ Estimable + ✓ Small + ✓ Testable +``` + +--- + +## Sprint Metrics + +Track sprint health and team performance. + +### Key Metrics + +| Metric | Formula | Target | +|--------|---------|--------| +| Velocity | Points completed / sprint | Stable ±10% | +| Commitment Reliability | Completed / Committed | >85% | +| Scope Change | Points added or removed mid-sprint | <10% | +| Carryover | Points not completed | <15% | + +### Velocity Tracking + +``` +Sprint 1: 25 points +Sprint 2: 28 points +Sprint 3: 30 points +Sprint 4: 32 points +Sprint 5: 29 points +------------------------ +Average Velocity: 28.8 points +Trend: Stable + +Planning: Commit to 24-26 points +``` + +### Definition of Done + +Story is complete when: + +- [ ] Code complete and peer reviewed +- [ ] Unit tests written and passing +- [ ] Acceptance criteria verified +- [ ] Documentation updated +- [ ] Deployed to staging environment +- [ ] Product Owner accepted +- [ ] No critical bugs remaining diff --git a/skills/agile-product-owner/_meta.json b/skills/agile-product-owner/_meta.json new file mode 100644 index 00000000..a2465229 --- /dev/null +++ b/skills/agile-product-owner/_meta.json @@ -0,0 +1,22 @@ +{ + "owner": "alirezarezvani", + "slug": "agile-product-owner", + "displayName": "Agile Product Owner", + "latest": { + "version": "2.1.1", + "publishedAt": 1773070352360, + "commit": "https://github.com/openclaw/skills/commit/c933c714efb0623916758946527dd266a712ee2b" + }, + "history": [ + { + "version": "1.0.0", + "publishedAt": 1770402562930, + "commit": "https://github.com/openclaw/skills/commit/64f88db738907b6a2c013929865bbd89ad000a5c" + }, + { + "version": "0.1.0", + "publishedAt": 1770022297220, + "commit": "https://github.com/clawdbot/skills/commit/5a797eb236d0e6bca9d3e23ff7113b2a402bbf86" + } + ] +} diff --git a/skills/agile-product-owner/references/sprint-planning-guide.md b/skills/agile-product-owner/references/sprint-planning-guide.md new file mode 100644 index 00000000..b21023bf --- /dev/null +++ b/skills/agile-product-owner/references/sprint-planning-guide.md @@ -0,0 +1,324 @@ +# Sprint Planning Guide + +Sprint planning workflows, capacity calculation, and backlog management. + +--- + +## Table of Contents + +- [Sprint Planning Workflow](#sprint-planning-workflow) +- [Capacity Planning](#capacity-planning) +- [Backlog Prioritization](#backlog-prioritization) +- [Sprint Ceremonies](#sprint-ceremonies) +- [Metrics and Tracking](#metrics-and-tracking) + +--- + +## Sprint Planning Workflow + +### Pre-Planning (1-2 Days Before) + +1. Review and refine backlog items for upcoming sprint +2. Ensure top items have acceptance criteria +3. Validate story point estimates with team +4. Identify dependencies between stories +5. Confirm team availability for sprint +6. **Validation:** Top 1.5x capacity of stories are refined and estimated + +### Sprint Planning Meeting + +**Duration:** 2 hours for 2-week sprint + +**Agenda:** + +| Time | Activity | Participants | +|------|----------|--------------| +| 0:00-0:15 | Review sprint goal and priorities | PO presents | +| 0:15-0:45 | Discuss top backlog items | Team asks questions | +| 0:45-1:15 | Team selects stories for sprint | Team decides | +| 1:15-1:45 | Break down stories into tasks | Team collaborates | +| 1:45-2:00 | Confirm commitment and identify risks | All | + +### Planning Checklist + +**Before Planning:** +- [ ] Backlog groomed with top items refined +- [ ] Previous sprint retrospective actions reviewed +- [ ] Team capacity calculated +- [ ] Dependencies identified +- [ ] Sprint goal drafted + +**During Planning:** +- [ ] Sprint goal agreed +- [ ] Stories selected fit within capacity +- [ ] Acceptance criteria reviewed for each story +- [ ] Tasks identified for complex stories +- [ ] Risks and blockers discussed + +**After Planning:** +- [ ] Sprint backlog visible to all +- [ ] Sprint goal communicated +- [ ] Calendar blocked for ceremonies +- [ ] Dependencies communicated to other teams + +--- + +## Capacity Planning + +### Team Capacity Calculation + +``` +Sprint Capacity = (Team Members × Sprint Days × Hours/Day × Focus Factor) + ÷ Hours per Story Point + +Simplified Version: +Sprint Capacity = Average Velocity × Availability Factor +``` + +### Availability Factors + +| Scenario | Factor | Example | +|----------|--------|---------| +| Full sprint, no PTO | 1.0 | 30 points if velocity = 30 | +| 1 team member out 50% | 0.9 | 27 points | +| Holiday during sprint | 0.8 | 24 points | +| Multiple team members out | 0.7 | 21 points | +| Major release/on-call | 0.75 | 22-23 points | + +### Capacity Buffer Rules + +| Commitment Level | % of Velocity | Purpose | +|------------------|---------------|---------| +| Committed | 80-85% | High confidence delivery | +| Stretch | 10-15% | Optional if things go well | +| Buffer | 5-10% | Unplanned work, bugs | + +### Sprint Loading Example + +``` +Team Velocity: 30 points/sprint +Availability: 90% (one team member partially out) +Adjusted Velocity: 27 points + +Sprint Loading: +- Committed work: 23 points (85% of 27) +- Stretch goals: 4 points (15% of 27) +- Buffer: Remaining capacity for bugs/support + +Story Selection: +[H] US-001: User dashboard (5 pts) ← Committed +[H] US-002: Export feature (3 pts) ← Committed +[H] US-003: Search filter (5 pts) ← Committed +[M] US-004: Settings page (5 pts) ← Committed +[M] US-005: Help tooltips (3 pts) ← Committed +[L] US-006: Theme options (2 pts) ← Committed +------------------------ +Committed Total: 23 points + +[L] US-007: Sort options (2 pts) ← Stretch +[L] US-008: Print view (2 pts) ← Stretch +------------------------ +Stretch Total: 4 points +``` + +--- + +## Backlog Prioritization + +### Priority Framework + +| Priority | Definition | SLA | +|----------|------------|-----| +| Critical | Blocking users, security, data loss | Immediate | +| High | Core functionality, key user needs | This sprint | +| Medium | Improvements, enhancements | Next 2-3 sprints | +| Low | Nice-to-have, minor improvements | Backlog | + +### Prioritization Factors + +| Factor | Weight | Questions | +|--------|--------|-----------| +| Business Value | 40% | Revenue impact? User demand? Strategic? | +| User Impact | 30% | How many users? How often used? | +| Risk/Dependencies | 15% | Technical risk? External dependencies? | +| Effort | 15% | Size? Complexity? Uncertainty? | + +### WSJF (Weighted Shortest Job First) + +For larger items, use SAFe's WSJF: + +``` +WSJF = Cost of Delay / Job Duration + +Cost of Delay = User Value + Time Criticality + Risk Reduction + +Scale: 1, 2, 3, 5, 8, 13, 20 + +Example: +Feature A: CoD = 13, Duration = 5 → WSJF = 2.6 +Feature B: CoD = 8, Duration = 2 → WSJF = 4.0 ← Higher priority +``` + +### Backlog Organization + +| Section | Content | Review Frequency | +|---------|---------|------------------| +| Sprint Backlog | Committed for current sprint | Daily | +| Ready | Refined, estimated, prioritized | Each planning | +| Grooming | Needs refinement | Weekly | +| Icebox | Future consideration | Monthly | +| Archive | Completed or obsolete | Quarterly | + +--- + +## Sprint Ceremonies + +### Daily Standup + +**Duration:** 15 minutes max +**Format:** Each team member answers: + +1. What did I complete yesterday? +2. What will I work on today? +3. What blockers do I have? + +**Product Owner Role:** +- Listen for blockers needing PO action +- Answer clarifying questions +- Note scope concerns for offline discussion +- Update stakeholders on progress + +### Backlog Refinement (Grooming) + +**Duration:** 1-2 hours per week +**Timing:** Mid-sprint + +**Agenda:** + +| Time | Activity | +|------|----------| +| 0:00-0:15 | Review upcoming priorities | +| 0:15-0:45 | Detail acceptance criteria for top items | +| 0:45-1:15 | Estimate new stories | +| 1:15-1:30 | Split large stories | + +**Readiness Criteria:** +- [ ] Clear user story format (As a... I want... So that...) +- [ ] Acceptance criteria defined (Given-When-Then) +- [ ] Story point estimate agreed +- [ ] Dependencies identified +- [ ] Fits in one sprint (≤8 points) + +### Sprint Review (Demo) + +**Duration:** 1 hour for 2-week sprint + +**Agenda:** + +| Time | Activity | Lead | +|------|----------|------| +| 0:00-0:05 | Sprint goal recap | PO | +| 0:05-0:40 | Demo completed work | Team | +| 0:40-0:50 | Stakeholder feedback | Stakeholders | +| 0:50-1:00 | Roadmap update | PO | + +**Demo Checklist:** +- [ ] Only demo completed (done-done) stories +- [ ] Use production or production-like environment +- [ ] Show user perspective, not technical details +- [ ] Collect feedback for backlog items +- [ ] Thank team for accomplishments + +### Sprint Retrospective + +**Duration:** 1.5 hours for 2-week sprint + +**Format Options:** + +| Format | Structure | +|--------|-----------| +| Start-Stop-Continue | What to begin, end, keep doing | +| 4Ls | Liked, Learned, Lacked, Longed for | +| Sailboat | Wind (helpers), Anchors (blockers), Rocks (risks) | +| Mad-Sad-Glad | Emotional state about sprint events | + +**Action Items:** +- Maximum 2-3 improvement actions per retro +- Assign owner and due date +- Review previous actions at start of next retro + +--- + +## Metrics and Tracking + +### Sprint Metrics + +| Metric | Formula | Target | +|--------|---------|--------| +| Velocity | Points completed / sprint | Stable ±10% | +| Commitment Reliability | Completed / Committed | >85% | +| Scope Change | Points added or removed | <10% | +| Carryover | Points not completed | <15% | +| Bug Ratio | Bug points / Total points | <20% | + +### Velocity Tracking + +``` +Sprint Velocity Trend: +Sprint 1: 25 points +Sprint 2: 28 points +Sprint 3: 30 points +Sprint 4: 32 points +Sprint 5: 29 points +------------------------ +Average: 28.8 points +Trend: Stable (±10%) + +Planning Recommendation: Plan for 26-29 points committed +``` + +### Burndown Chart + +Track progress within sprint: + +``` +Day Ideal Actual Status +--- ----- ------ ------ + 0 30 30 On track + 2 24 26 Slightly behind + 4 18 20 Behind + 6 12 14 Recovering + 8 6 6 On track +10 0 2 Minor carryover +``` + +**Burndown Patterns:** + +| Pattern | Meaning | Action | +|---------|---------|--------| +| Flat start | No progress early | Check blockers | +| Late drop | Last-minute completion | Improve WIP limits | +| Scope increase | Line moves up | Address scope creep | +| Early completion | Done before sprint end | Pull stretch items | + +### Definition of Done + +Story is complete when: + +- [ ] Code complete and reviewed +- [ ] Unit tests written and passing +- [ ] Integration tests passing +- [ ] Acceptance criteria verified +- [ ] Documentation updated +- [ ] Deployed to staging +- [ ] PO accepted +- [ ] No critical bugs + +### Release Metrics + +| Metric | Definition | Target | +|--------|------------|--------| +| Lead Time | Idea to production | <2 sprints | +| Cycle Time | Development start to done | <1 sprint | +| Throughput | Stories completed/sprint | Increasing | +| Defect Escape | Bugs found in production | Decreasing | diff --git a/skills/agile-product-owner/references/user-story-templates.md b/skills/agile-product-owner/references/user-story-templates.md new file mode 100644 index 00000000..ed2517fc --- /dev/null +++ b/skills/agile-product-owner/references/user-story-templates.md @@ -0,0 +1,291 @@ +# User Story Templates + +Standard templates, acceptance criteria patterns, and INVEST validation for user stories. + +--- + +## Table of Contents + +- [Story Templates](#story-templates) +- [Acceptance Criteria Patterns](#acceptance-criteria-patterns) +- [INVEST Criteria](#invest-criteria) +- [Story Point Estimation](#story-point-estimation) +- [Common Antipatterns](#common-antipatterns) + +--- + +## Story Templates + +### Standard User Story Format + +``` +As a [persona], +I want to [action/capability], +So that [benefit/value]. +``` + +### Template by Story Type + +**Feature Story:** +``` +As a [persona], +I want to [perform action] +So that [I achieve benefit]. + +Example: +As a marketing manager, +I want to export campaign reports to PDF +So that I can share results with stakeholders who don't have system access. +``` + +**Improvement Story:** +``` +As a [persona], +I need [capability/improvement] +To [achieve goal more effectively]. + +Example: +As a sales rep, +I need faster search results +To find customer records without interrupting calls. +``` + +**Bug Fix Story:** +``` +As a [persona], +I expect [correct behavior] +When [specific condition]. + +Example: +As a user, +I expect my session to remain active +When navigating between dashboard tabs. +``` + +**Integration Story:** +``` +As a [persona], +I want to [integrate/connect with system] +So that [workflow improvement]. + +Example: +As an admin, +I want to sync user data with our LDAP server +So that employees are automatically provisioned. +``` + +**Enabler Story (Technical):** +``` +As a developer, +I need to [technical requirement] +To enable [user-facing capability]. + +Example: +As a developer, +I need to implement caching layer +To enable sub-second dashboard load times. +``` + +### Persona Library + +| Persona | Typical Needs | Context | +|---------|--------------|---------| +| End User | Efficiency, simplicity, reliability | Daily core feature usage | +| Administrator | Control, visibility, security | System management | +| Power User | Automation, customization, shortcuts | Expert workflows | +| New User | Guidance, learning, safety | Onboarding experience | +| Manager | Reporting, oversight, delegation | Team coordination | +| External User | Access, security, documentation | Customer/partner usage | + +--- + +## Acceptance Criteria Patterns + +### Given-When-Then (Gherkin) + +Preferred format for testable acceptance criteria: + +``` +Given [precondition/context], +When [action/trigger], +Then [expected outcome]. +``` + +**Examples:** + +``` +Given the user is logged in with valid credentials, +When they click the "Export" button, +Then a PDF download starts within 2 seconds. + +Given the user has entered invalid email format, +When they submit the registration form, +Then an inline error message displays "Please enter a valid email address." + +Given the daily sync job has not run in 24 hours, +When the scheduler triggers at midnight, +Then all pending records are synchronized and logged. +``` + +### Should/Must/Can Patterns + +**Should (Expected Behavior):** +``` +Should [behavior] when [condition]. + +Example: +Should display loading spinner when API call exceeds 500ms. +``` + +**Must (Hard Requirement):** +``` +Must [requirement] to [achieve outcome]. + +Example: +Must encrypt all data at rest to meet compliance requirements. +``` + +**Can (Capability):** +``` +Can [capability] without [negative outcome]. + +Example: +Can undo last action without losing other changes. +``` + +### Acceptance Criteria Checklist + +Each story should have acceptance criteria covering: + +| Category | Example Criterion | +|----------|-------------------| +| Happy Path | Given valid input, When submitted, Then success message displayed | +| Validation | Should reject input when required field is empty | +| Error Handling | Must show user-friendly message when API fails | +| Performance | Should complete operation within 2 seconds | +| Accessibility | Must be navigable via keyboard only | +| Security | Should not expose sensitive data in URL parameters | + +### Minimum Acceptance Criteria Count + +| Story Size (Points) | Minimum AC Count | +|--------------------|------------------| +| 1-2 | 3-4 | +| 3-5 | 4-6 | +| 8 | 5-8 | +| 13+ | Split the story | + +--- + +## INVEST Criteria + +### INVEST Validation Checklist + +| Criterion | Question | Pass If... | +|-----------|----------|------------| +| **I**ndependent | Can this story be developed without depending on another story? | No blocking dependencies on uncommitted work | +| **N**egotiable | Is the implementation approach flexible? | Multiple ways to deliver the value | +| **V**aluable | Does this deliver value to users or business? | Clear benefit statement in "so that" | +| **E**stimable | Can the team estimate this story? | Understood well enough to size | +| **S**mall | Can this be completed in one sprint? | ≤8 story points typically | +| **T**estable | Can we verify this story is done? | Clear, measurable acceptance criteria | + +### INVEST Failure Patterns + +| Criterion | Red Flag | Fix | +|-----------|----------|-----| +| Independent | "After story X is done..." | Combine stories or resequence | +| Negotiable | Specific implementation in story | Focus on outcome, not solution | +| Valuable | No "so that" clause | Add benefit statement | +| Estimable | Team says "no idea" | Spike first, then story | +| Small | >8 points | Split into smaller stories | +| Testable | "System should be better" | Add measurable criteria | + +### Story Splitting Techniques + +When stories are too large (>8 points), split using: + +| Technique | Example | +|-----------|---------| +| By workflow step | "Create order" → "Add items" + "Apply discount" + "Submit order" | +| By persona | "User dashboard" → "Admin dashboard" + "Member dashboard" | +| By data type | "Import data" → "Import CSV" + "Import Excel" | +| By operation | "Manage users" → "Add user" + "Edit user" + "Delete user" | +| By platform | "Mobile support" → "iOS support" + "Android support" | +| Happy path first | "Full feature" → "Basic feature" + "Error handling" + "Edge cases" | + +--- + +## Story Point Estimation + +### Fibonacci Scale Reference + +| Points | Complexity | Example | +|--------|------------|---------| +| 1 | Trivial | Fix typo, change label | +| 2 | Simple | Add field, simple validation | +| 3 | Small | New form, basic CRUD operation | +| 5 | Medium | Feature with multiple components | +| 8 | Large | Complex feature, multiple integrations | +| 13 | Very Large | Consider splitting | +| 21+ | Epic | Must split | + +### Estimation Factors + +| Factor | Low Complexity | High Complexity | +|--------|---------------|-----------------| +| Unknowns | Well understood | Many unknowns | +| Dependencies | None | Multiple systems | +| Testing | Simple unit tests | Complex integration tests | +| Data | Simple structure | Complex transformations | +| UI | Minor changes | New components | + +### Velocity Calculation + +``` +Velocity = Total points completed / Number of sprints + +Example: +Sprint 1: 28 points +Sprint 2: 32 points +Sprint 3: 30 points +Average Velocity: (28 + 32 + 30) / 3 = 30 points/sprint + +Sprint Capacity Planning: +- Committed: 80-90% of velocity (24-27 points) +- Stretch goals: 10-20% additional (3-6 points) +``` + +--- + +## Common Antipatterns + +### Story Antipatterns + +| Antipattern | Example | Fix | +|-------------|---------|-----| +| Solution story | "Implement React component" | "Display user profile information" | +| Compound story | "Create, edit, and delete users" | Split into three stories | +| Missing persona | "The system will..." | "As an admin, I want to..." | +| No benefit | "I want to see a button" | Add "so that [benefit]" | +| Too vague | "Improve performance" | "Reduce page load to <2 seconds" | +| Technical jargon | "Implement Redis caching" | "Enable instant search results" | + +### Acceptance Criteria Antipatterns + +| Antipattern | Example | Fix | +|-------------|---------|-----| +| Too vague | "Works correctly" | Specific Given-When-Then | +| Implementation details | "Use PostgreSQL query" | Focus on outcome | +| Missing unhappy path | Only success scenario | Add error cases | +| Untestable | "User is happy" | Measurable behavior | +| Too many | 15+ criteria | Split the story | + +### Sprint Planning Antipatterns + +| Antipattern | Impact | Fix | +|-------------|--------|-----| +| 100% capacity | No buffer for unknowns | Plan 80-85% | +| All large stories | Risk of incomplete sprint | Mix sizes | +| No dependencies mapped | Blocked work | Identify dependencies upfront | +| Stretch = overflow | Hiding overcommitment | Stretch should be optional | diff --git a/skills/agile-product-owner/scripts/user_story_generator.py b/skills/agile-product-owner/scripts/user_story_generator.py new file mode 100644 index 00000000..29471f84 --- /dev/null +++ b/skills/agile-product-owner/scripts/user_story_generator.py @@ -0,0 +1,387 @@ +#!/usr/bin/env python3 +""" +User Story Generator with INVEST Criteria +Creates well-formed user stories with acceptance criteria +""" + +import json +from typing import Dict, List, Tuple + +class UserStoryGenerator: + """Generate INVEST-compliant user stories""" + + def __init__(self): + self.personas = { + 'end_user': { + 'name': 'End User', + 'needs': ['efficiency', 'simplicity', 'reliability', 'speed'], + 'context': 'daily usage of core features' + }, + 'admin': { + 'name': 'Administrator', + 'needs': ['control', 'visibility', 'security', 'configuration'], + 'context': 'system management and oversight' + }, + 'power_user': { + 'name': 'Power User', + 'needs': ['advanced features', 'automation', 'customization', 'shortcuts'], + 'context': 'expert usage and workflow optimization' + }, + 'new_user': { + 'name': 'New User', + 'needs': ['guidance', 'learning', 'safety', 'clarity'], + 'context': 'first-time experience and onboarding' + } + } + + self.story_templates = { + 'feature': "As a {persona}, I want to {action} so that {benefit}", + 'improvement': "As a {persona}, I need {capability} to {achieve_goal}", + 'fix': "As a {persona}, I expect {behavior} when {condition}", + 'integration': "As a {persona}, I want to {integrate} so that {workflow}" + } + + self.acceptance_criteria_patterns = [ + "Given {precondition}, When {action}, Then {outcome}", + "Should {behavior} when {condition}", + "Must {requirement} to {achieve}", + "Can {capability} without {negative_outcome}" + ] + + def generate_epic_stories(self, epic: Dict) -> List[Dict]: + """Break down epic into user stories""" + stories = [] + + # Analyze epic for key components + epic_name = epic.get('name', 'Feature') + epic_description = epic.get('description', '') + personas = epic.get('personas', ['end_user']) + scope = epic.get('scope', []) + + # Generate stories for each persona and scope item + for persona in personas: + for i, scope_item in enumerate(scope): + story = self.generate_story( + persona=persona, + feature=scope_item, + epic=epic_name, + index=i+1 + ) + stories.append(story) + + # Add enabler stories (technical, infrastructure) + if epic.get('technical_requirements'): + for req in epic['technical_requirements']: + enabler = self.generate_enabler_story(req, epic_name) + stories.append(enabler) + + return stories + + def generate_story(self, persona: str, feature: str, epic: str, index: int) -> Dict: + """Generate a single user story""" + + persona_data = self.personas.get(persona, self.personas['end_user']) + + # Create story + story = { + 'id': f"{epic[:3].upper()}-{index:03d}", + 'type': 'story', + 'title': self._generate_title(feature), + 'narrative': self._generate_narrative(persona_data, feature), + 'acceptance_criteria': self._generate_acceptance_criteria(feature), + 'estimation': self._estimate_complexity(feature), + 'priority': self._determine_priority(persona, feature), + 'dependencies': [], + 'invest_check': self._check_invest_criteria(feature) + } + + return story + + def generate_enabler_story(self, requirement: str, epic: str) -> Dict: + """Generate technical enabler story""" + + return { + 'id': f"{epic[:3].upper()}-E{len(requirement):02d}", + 'type': 'enabler', + 'title': f"Technical: {requirement}", + 'narrative': f"As a developer, I need to {requirement} to enable user features", + 'acceptance_criteria': [ + f"Technical requirement {requirement} is implemented", + "All tests pass", + "Documentation is updated", + "No regression in existing functionality" + ], + 'estimation': 5, # Default medium complexity + 'priority': 'high', + 'dependencies': [], + 'invest_check': { + 'independent': True, + 'negotiable': False, # Technical requirements often non-negotiable + 'valuable': True, + 'estimable': True, + 'small': True, + 'testable': True + } + } + + def _generate_title(self, feature: str) -> str: + """Generate concise story title""" + # Simplify feature description to title + words = feature.split()[:5] + return ' '.join(words).title() + + def _generate_narrative(self, persona: Dict, feature: str) -> str: + """Generate story narrative in standard format""" + + template = self.story_templates['feature'] + + action = self._extract_action(feature) + benefit = self._extract_benefit(feature, persona['needs']) + + return template.format( + persona=persona['name'], + action=action, + benefit=benefit + ) + + def _generate_acceptance_criteria(self, feature: str) -> List[str]: + """Generate acceptance criteria""" + + criteria = [] + + # Happy path + criteria.append(f"Given user has access, When they {self._extract_action(feature)}, Then {self._extract_outcome(feature)}") + + # Validation + criteria.append(f"Should validate input before processing") + + # Error handling + criteria.append(f"Must show clear error message when action fails") + + # Performance + criteria.append(f"Should complete within 2 seconds") + + # Accessibility + criteria.append(f"Must be accessible via keyboard navigation") + + return criteria + + def _extract_action(self, feature: str) -> str: + """Extract action from feature description""" + action_verbs = ['create', 'view', 'edit', 'delete', 'share', 'export', 'import', 'configure', 'search', 'filter'] + + feature_lower = feature.lower() + for verb in action_verbs: + if verb in feature_lower: + return feature_lower + + return f"use {feature.lower()}" + + def _extract_benefit(self, feature: str, needs: List[str]) -> str: + """Extract benefit based on feature and persona needs""" + + feature_lower = feature.lower() + + if 'save' in feature_lower or 'quick' in feature_lower: + return "I can save time and work more efficiently" + elif 'share' in feature_lower or 'collab' in feature_lower: + return "I can collaborate with my team effectively" + elif 'report' in feature_lower or 'analyt' in feature_lower: + return "I can make data-driven decisions" + elif 'automat' in feature_lower: + return "I can reduce manual work and errors" + else: + return f"I can achieve my goals related to {needs[0]}" + + def _extract_outcome(self, feature: str) -> str: + """Extract expected outcome""" + return f"the {feature.lower()} is successfully completed" + + def _estimate_complexity(self, feature: str) -> int: + """Estimate story points based on complexity indicators""" + + feature_lower = feature.lower() + + # Complexity indicators + complexity = 3 # Base complexity + + if any(word in feature_lower for word in ['simple', 'basic', 'view', 'display']): + complexity = 1 + elif any(word in feature_lower for word in ['create', 'edit', 'update']): + complexity = 3 + elif any(word in feature_lower for word in ['complex', 'advanced', 'integrate', 'migrate']): + complexity = 8 + elif any(word in feature_lower for word in ['redesign', 'refactor', 'architect']): + complexity = 13 + + return complexity + + def _determine_priority(self, persona: str, feature: str) -> str: + """Determine story priority""" + + feature_lower = feature.lower() + + # Critical features + if any(word in feature_lower for word in ['security', 'fix', 'critical', 'broken']): + return 'critical' + + # High priority for primary personas + if persona in ['end_user', 'admin']: + if any(word in feature_lower for word in ['core', 'essential', 'primary']): + return 'high' + + # Medium for improvements + if any(word in feature_lower for word in ['improve', 'enhance', 'optimize']): + return 'medium' + + # Low for nice-to-haves + return 'low' + + def _check_invest_criteria(self, feature: str) -> Dict[str, bool]: + """Check INVEST criteria compliance""" + + return { + 'independent': not any(word in feature.lower() for word in ['after', 'depends', 'requires']), + 'negotiable': True, # Most features can be negotiated + 'valuable': True, # Assume value if it made it to backlog + 'estimable': len(feature.split()) < 20, # Can estimate if not too vague + 'small': self._estimate_complexity(feature) <= 8, # 8 points or less + 'testable': not any(word in feature.lower() for word in ['maybe', 'possibly', 'somehow']) + } + + def generate_sprint_stories(self, capacity: int, backlog: List[Dict]) -> Dict: + """Generate stories for a sprint based on capacity""" + + sprint = { + 'capacity': capacity, + 'committed': [], + 'stretch': [], + 'total_points': 0, + 'utilization': 0 + } + + # Sort backlog by priority and size + sorted_backlog = sorted( + backlog, + key=lambda x: ( + {'critical': 0, 'high': 1, 'medium': 2, 'low': 3}[x['priority']], + x['estimation'] + ) + ) + + # Fill sprint + for story in sorted_backlog: + if sprint['total_points'] + story['estimation'] <= capacity: + sprint['committed'].append(story) + sprint['total_points'] += story['estimation'] + elif sprint['total_points'] + story['estimation'] <= capacity * 1.2: + sprint['stretch'].append(story) + + sprint['utilization'] = round((sprint['total_points'] / capacity) * 100, 1) + + return sprint + + def format_story_output(self, story: Dict) -> str: + """Format story for display""" + + output = [] + output.append(f"USER STORY: {story['id']}") + output.append("=" * 40) + output.append(f"Title: {story['title']}") + output.append(f"Type: {story['type']}") + output.append(f"Priority: {story['priority'].upper()}") + output.append(f"Points: {story['estimation']}") + output.append("") + output.append("Story:") + output.append(story['narrative']) + output.append("") + output.append("Acceptance Criteria:") + for i, criterion in enumerate(story['acceptance_criteria'], 1): + output.append(f" {i}. {criterion}") + output.append("") + output.append("INVEST Checklist:") + for criterion, passed in story['invest_check'].items(): + status = "✓" if passed else "✗" + output.append(f" {status} {criterion.capitalize()}") + + return "\n".join(output) + +def create_sample_epic(): + """Create a sample epic for testing""" + return { + 'name': 'User Dashboard', + 'description': 'Create a comprehensive dashboard for users to view their data', + 'personas': ['end_user', 'power_user'], + 'scope': [ + 'View key metrics and KPIs', + 'Customize dashboard layout', + 'Export dashboard data', + 'Share dashboard with team members', + 'Set up automated reports' + ], + 'technical_requirements': [ + 'Implement caching for performance', + 'Set up real-time data pipeline' + ] + } + +def main(): + import sys + + generator = UserStoryGenerator() + + if len(sys.argv) > 1 and sys.argv[1] == 'sprint': + # Generate sprint planning + capacity = int(sys.argv[2]) if len(sys.argv) > 2 else 30 + + # Create sample backlog + epic = create_sample_epic() + backlog = generator.generate_epic_stories(epic) + + # Plan sprint + sprint = generator.generate_sprint_stories(capacity, backlog) + + print("=" * 60) + print("SPRINT PLANNING") + print("=" * 60) + print(f"Sprint Capacity: {sprint['capacity']} points") + print(f"Committed: {sprint['total_points']} points ({sprint['utilization']}%)") + print(f"Stories: {len(sprint['committed'])} committed + {len(sprint['stretch'])} stretch") + print("\n📋 COMMITTED STORIES:\n") + + for story in sprint['committed']: + print(f" [{story['priority'][:1].upper()}] {story['id']}: {story['title']} ({story['estimation']}pts)") + + if sprint['stretch']: + print("\n🎯 STRETCH GOALS:\n") + for story in sprint['stretch']: + print(f" [{story['priority'][:1].upper()}] {story['id']}: {story['title']} ({story['estimation']}pts)") + + else: + # Generate stories for epic + epic = create_sample_epic() + stories = generator.generate_epic_stories(epic) + + print(f"Generated {len(stories)} stories from epic: {epic['name']}\n") + + # Display first 3 stories in detail + for story in stories[:3]: + print(generator.format_story_output(story)) + print("\n") + + # Summary of all stories + print("=" * 60) + print("BACKLOG SUMMARY") + print("=" * 60) + total_points = sum(s['estimation'] for s in stories) + print(f"Total Stories: {len(stories)}") + print(f"Total Points: {total_points}") + print(f"Average Size: {total_points/len(stories):.1f} points") + print("\nPriority Breakdown:") + for priority in ['critical', 'high', 'medium', 'low']: + count = len([s for s in stories if s['priority'] == priority]) + if count > 0: + print(f" {priority.capitalize()}: {count} stories") + +if __name__ == "__main__": + main() diff --git a/skills/api-design-reviewer/SKILL.md b/skills/api-design-reviewer/SKILL.md new file mode 100644 index 00000000..4fafc63a --- /dev/null +++ b/skills/api-design-reviewer/SKILL.md @@ -0,0 +1,421 @@ +--- +name: "api-design-reviewer" +description: "API Design Reviewer" +--- + +# API Design Reviewer + +**Tier:** POWERFUL +**Category:** Engineering / Architecture +**Maintainer:** Claude Skills Team + +## Overview + +The API Design Reviewer skill provides comprehensive analysis and review of API designs, focusing on REST conventions, best practices, and industry standards. This skill helps engineering teams build consistent, maintainable, and well-designed APIs through automated linting, breaking change detection, and design scorecards. + +## Core Capabilities + +### 1. API Linting and Convention Analysis +- **Resource Naming Conventions**: Enforces kebab-case for resources, camelCase for fields +- **HTTP Method Usage**: Validates proper use of GET, POST, PUT, PATCH, DELETE +- **URL Structure**: Analyzes endpoint patterns for consistency and RESTful design +- **Status Code Compliance**: Ensures appropriate HTTP status codes are used +- **Error Response Formats**: Validates consistent error response structures +- **Documentation Coverage**: Checks for missing descriptions and documentation gaps + +### 2. Breaking Change Detection +- **Endpoint Removal**: Detects removed or deprecated endpoints +- **Response Shape Changes**: Identifies modifications to response structures +- **Field Removal**: Tracks removed or renamed fields in API responses +- **Type Changes**: Catches field type modifications that could break clients +- **Required Field Additions**: Flags new required fields that could break existing integrations +- **Status Code Changes**: Detects changes to expected status codes + +### 3. API Design Scoring and Assessment +- **Consistency Analysis** (30%): Evaluates naming conventions, response patterns, and structural consistency +- **Documentation Quality** (20%): Assesses completeness and clarity of API documentation +- **Security Implementation** (20%): Reviews authentication, authorization, and security headers +- **Usability Design** (15%): Analyzes ease of use, discoverability, and developer experience +- **Performance Patterns** (15%): Evaluates caching, pagination, and efficiency patterns + +## REST Design Principles + +### Resource Naming Conventions +``` +✅ Good Examples: +- /api/v1/users +- /api/v1/user-profiles +- /api/v1/orders/123/line-items + +❌ Bad Examples: +- /api/v1/getUsers +- /api/v1/user_profiles +- /api/v1/orders/123/lineItems +``` + +### HTTP Method Usage +- **GET**: Retrieve resources (safe, idempotent) +- **POST**: Create new resources (not idempotent) +- **PUT**: Replace entire resources (idempotent) +- **PATCH**: Partial resource updates (not necessarily idempotent) +- **DELETE**: Remove resources (idempotent) + +### URL Structure Best Practices +``` +Collection Resources: /api/v1/users +Individual Resources: /api/v1/users/123 +Nested Resources: /api/v1/users/123/orders +Actions: /api/v1/users/123/activate (POST) +Filtering: /api/v1/users?status=active&role=admin +``` + +## Versioning Strategies + +### 1. URL Versioning (Recommended) +``` +/api/v1/users +/api/v2/users +``` +**Pros**: Clear, explicit, easy to route +**Cons**: URL proliferation, caching complexity + +### 2. Header Versioning +``` +GET /api/users +Accept: application/vnd.api+json;version=1 +``` +**Pros**: Clean URLs, content negotiation +**Cons**: Less visible, harder to test manually + +### 3. Media Type Versioning +``` +GET /api/users +Accept: application/vnd.myapi.v1+json +``` +**Pros**: RESTful, supports multiple representations +**Cons**: Complex, harder to implement + +### 4. Query Parameter Versioning +``` +/api/users?version=1 +``` +**Pros**: Simple to implement +**Cons**: Not RESTful, can be ignored + +## Pagination Patterns + +### Offset-Based Pagination +```json +{ + "data": [...], + "pagination": { + "offset": 20, + "limit": 10, + "total": 150, + "hasMore": true + } +} +``` + +### Cursor-Based Pagination +```json +{ + "data": [...], + "pagination": { + "nextCursor": "eyJpZCI6MTIzfQ==", + "hasMore": true + } +} +``` + +### Page-Based Pagination +```json +{ + "data": [...], + "pagination": { + "page": 3, + "pageSize": 10, + "totalPages": 15, + "totalItems": 150 + } +} +``` + +## Error Response Formats + +### Standard Error Structure +```json +{ + "error": { + "code": "VALIDATION_ERROR", + "message": "The request contains invalid parameters", + "details": [ + { + "field": "email", + "code": "INVALID_FORMAT", + "message": "Email address is not valid" + } + ], + "requestId": "req-123456", + "timestamp": "2024-02-16T13:00:00Z" + } +} +``` + +### HTTP Status Code Usage +- **400 Bad Request**: Invalid request syntax or parameters +- **401 Unauthorized**: Authentication required +- **403 Forbidden**: Access denied (authenticated but not authorized) +- **404 Not Found**: Resource not found +- **409 Conflict**: Resource conflict (duplicate, version mismatch) +- **422 Unprocessable Entity**: Valid syntax but semantic errors +- **429 Too Many Requests**: Rate limit exceeded +- **500 Internal Server Error**: Unexpected server error + +## Authentication and Authorization Patterns + +### Bearer Token Authentication +``` +Authorization: Bearer +``` + +### API Key Authentication +``` +X-API-Key: +Authorization: Api-Key +``` + +### OAuth 2.0 Flow +``` +Authorization: Bearer +``` + +### Role-Based Access Control (RBAC) +```json +{ + "user": { + "id": "123", + "roles": ["admin", "editor"], + "permissions": ["read:users", "write:orders"] + } +} +``` + +## Rate Limiting Implementation + +### Headers +``` +X-RateLimit-Limit: 1000 +X-RateLimit-Remaining: 999 +X-RateLimit-Reset: 1640995200 +``` + +### Response on Limit Exceeded +```json +{ + "error": { + "code": "RATE_LIMIT_EXCEEDED", + "message": "Too many requests", + "retryAfter": 3600 + } +} +``` + +## HATEOAS (Hypermedia as the Engine of Application State) + +### Example Implementation +```json +{ + "id": "123", + "name": "John Doe", + "email": "john@example.com", + "_links": { + "self": { "href": "/api/v1/users/123" }, + "orders": { "href": "/api/v1/users/123/orders" }, + "profile": { "href": "/api/v1/users/123/profile" }, + "deactivate": { + "href": "/api/v1/users/123/deactivate", + "method": "POST" + } + } +} +``` + +## Idempotency + +### Idempotent Methods +- **GET**: Always safe and idempotent +- **PUT**: Should be idempotent (replace entire resource) +- **DELETE**: Should be idempotent (same result) +- **PATCH**: May or may not be idempotent + +### Idempotency Keys +``` +POST /api/v1/payments +Idempotency-Key: 123e4567-e89b-12d3-a456-426614174000 +``` + +## Backward Compatibility Guidelines + +### Safe Changes (Non-Breaking) +- Adding optional fields to requests +- Adding fields to responses +- Adding new endpoints +- Making required fields optional +- Adding new enum values (with graceful handling) + +### Breaking Changes (Require Version Bump) +- Removing fields from responses +- Making optional fields required +- Changing field types +- Removing endpoints +- Changing URL structures +- Modifying error response formats + +## OpenAPI/Swagger Validation + +### Required Components +- **API Information**: Title, description, version +- **Server Information**: Base URLs and descriptions +- **Path Definitions**: All endpoints with methods +- **Parameter Definitions**: Query, path, header parameters +- **Request/Response Schemas**: Complete data models +- **Security Definitions**: Authentication schemes +- **Error Responses**: Standard error formats + +### Best Practices +- Use consistent naming conventions +- Provide detailed descriptions for all components +- Include examples for complex objects +- Define reusable components and schemas +- Validate against OpenAPI specification + +## Performance Considerations + +### Caching Strategies +``` +Cache-Control: public, max-age=3600 +ETag: "123456789" +Last-Modified: Wed, 21 Oct 2015 07:28:00 GMT +``` + +### Efficient Data Transfer +- Use appropriate HTTP methods +- Implement field selection (`?fields=id,name,email`) +- Support compression (gzip) +- Implement efficient pagination +- Use ETags for conditional requests + +### Resource Optimization +- Avoid N+1 queries +- Implement batch operations +- Use async processing for heavy operations +- Support partial updates (PATCH) + +## Security Best Practices + +### Input Validation +- Validate all input parameters +- Sanitize user data +- Use parameterized queries +- Implement request size limits + +### Authentication Security +- Use HTTPS everywhere +- Implement secure token storage +- Support token expiration and refresh +- Use strong authentication mechanisms + +### Authorization Controls +- Implement principle of least privilege +- Use resource-based permissions +- Support fine-grained access control +- Audit access patterns + +## Tools and Scripts + +### api_linter.py +Analyzes API specifications for compliance with REST conventions and best practices. + +**Features:** +- OpenAPI/Swagger spec validation +- Naming convention checks +- HTTP method usage validation +- Error format consistency +- Documentation completeness analysis + +### breaking_change_detector.py +Compares API specification versions to identify breaking changes. + +**Features:** +- Endpoint comparison +- Schema change detection +- Field removal/modification tracking +- Migration guide generation +- Impact severity assessment + +### api_scorecard.py +Provides comprehensive scoring of API design quality. + +**Features:** +- Multi-dimensional scoring +- Detailed improvement recommendations +- Letter grade assessment (A-F) +- Benchmark comparisons +- Progress tracking + +## Integration Examples + +### CI/CD Integration +```yaml +- name: "api-linting" + run: python scripts/api_linter.py openapi.json + +- name: "breaking-change-detection" + run: python scripts/breaking_change_detector.py openapi-v1.json openapi-v2.json + +- name: "api-scorecard" + run: python scripts/api_scorecard.py openapi.json +``` + +### Pre-commit Hooks +```bash +#!/bin/bash +python engineering/api-design-reviewer/scripts/api_linter.py api/openapi.json +if [ $? -ne 0 ]; then + echo "API linting failed. Please fix the issues before committing." + exit 1 +fi +``` + +## Best Practices Summary + +1. **Consistency First**: Maintain consistent naming, response formats, and patterns +2. **Documentation**: Provide comprehensive, up-to-date API documentation +3. **Versioning**: Plan for evolution with clear versioning strategies +4. **Error Handling**: Implement consistent, informative error responses +5. **Security**: Build security into every layer of the API +6. **Performance**: Design for scale and efficiency from the start +7. **Backward Compatibility**: Minimize breaking changes and provide migration paths +8. **Testing**: Implement comprehensive testing including contract testing +9. **Monitoring**: Add observability for API usage and performance +10. **Developer Experience**: Prioritize ease of use and clear documentation + +## Common Anti-Patterns to Avoid + +1. **Verb-based URLs**: Use nouns for resources, not actions +2. **Inconsistent Response Formats**: Maintain standard response structures +3. **Over-nesting**: Avoid deeply nested resource hierarchies +4. **Ignoring HTTP Status Codes**: Use appropriate status codes for different scenarios +5. **Poor Error Messages**: Provide actionable, specific error information +6. **Missing Pagination**: Always paginate list endpoints +7. **No Versioning Strategy**: Plan for API evolution from day one +8. **Exposing Internal Structure**: Design APIs for external consumption, not internal convenience +9. **Missing Rate Limiting**: Protect your API from abuse and overload +10. **Inadequate Testing**: Test all aspects including error cases and edge conditions + +## Conclusion + +The API Design Reviewer skill provides a comprehensive framework for building, reviewing, and maintaining high-quality REST APIs. By following these guidelines and using the provided tools, development teams can create APIs that are consistent, well-documented, secure, and maintainable. + +Regular use of the linting, breaking change detection, and scoring tools ensures continuous improvement and helps maintain API quality throughout the development lifecycle. \ No newline at end of file diff --git a/skills/api-design-reviewer/_meta.json b/skills/api-design-reviewer/_meta.json new file mode 100644 index 00000000..002b988c --- /dev/null +++ b/skills/api-design-reviewer/_meta.json @@ -0,0 +1,17 @@ +{ + "owner": "alirezarezvani", + "slug": "api-design-reviewer", + "displayName": "Api Design Reviewer", + "latest": { + "version": "2.1.1", + "publishedAt": 1773120995537, + "commit": "https://github.com/openclaw/skills/commit/1776f38e35fe6869ad8fdc19c621b96cbf0a5ab8" + }, + "history": [ + { + "version": "1.0.0", + "publishedAt": 1771261835499, + "commit": "https://github.com/openclaw/skills/commit/a22b421696c2b36082357ac42394da8041dc983c" + } + ] +} diff --git a/skills/api-design-reviewer/references/api_antipatterns.md b/skills/api-design-reviewer/references/api_antipatterns.md new file mode 100644 index 00000000..1e2bb99c --- /dev/null +++ b/skills/api-design-reviewer/references/api_antipatterns.md @@ -0,0 +1,680 @@ +# Common API Anti-Patterns and How to Avoid Them + +## Introduction + +This document outlines common anti-patterns in REST API design that can lead to poor developer experience, maintenance nightmares, and scalability issues. Each anti-pattern is accompanied by examples and recommended solutions. + +## 1. Verb-Based URLs (The RPC Trap) + +### Anti-Pattern +Using verbs in URLs instead of treating endpoints as resources. + +``` +❌ Bad Examples: +POST /api/getUsers +POST /api/createUser +GET /api/deleteUser/123 +POST /api/updateUserPassword +GET /api/calculateOrderTotal/456 +``` + +### Why It's Bad +- Violates REST principles +- Makes the API feel like RPC instead of REST +- HTTP methods lose their semantic meaning +- Reduces cacheability +- Harder to understand resource relationships + +### Solution +``` +✅ Good Examples: +GET /api/users # Get users +POST /api/users # Create user +DELETE /api/users/123 # Delete user +PATCH /api/users/123/password # Update password +GET /api/orders/456/total # Get order total +``` + +## 2. Inconsistent Naming Conventions + +### Anti-Pattern +Mixed naming conventions across the API. + +```json +❌ Bad Examples: +{ + "user_id": 123, // snake_case + "firstName": "John", // camelCase + "last-name": "Doe", // kebab-case + "EMAIL": "john@example.com", // UPPER_CASE + "IsActive": true // PascalCase +} +``` + +### Why It's Bad +- Confuses developers +- Increases cognitive load +- Makes code generation difficult +- Reduces API adoption + +### Solution +```json +✅ Choose one convention and stick to it (camelCase recommended): +{ + "userId": 123, + "firstName": "John", + "lastName": "Doe", + "email": "john@example.com", + "isActive": true +} +``` + +## 3. Ignoring HTTP Status Codes + +### Anti-Pattern +Always returning HTTP 200 regardless of the actual result. + +```json +❌ Bad Example: +HTTP/1.1 200 OK +{ + "status": "error", + "code": 404, + "message": "User not found" +} +``` + +### Why It's Bad +- Breaks HTTP semantics +- Prevents proper error handling by clients +- Breaks caching and proxies +- Makes monitoring and debugging harder + +### Solution +```json +✅ Good Example: +HTTP/1.1 404 Not Found +{ + "error": { + "code": "USER_NOT_FOUND", + "message": "User with ID 123 not found", + "requestId": "req-abc123" + } +} +``` + +## 4. Overly Complex Nested Resources + +### Anti-Pattern +Creating deeply nested URL structures that are hard to navigate. + +``` +❌ Bad Example: +/companies/123/departments/456/teams/789/members/012/projects/345/tasks/678/comments/901 +``` + +### Why It's Bad +- URLs become unwieldy +- Creates tight coupling between resources +- Makes independent resource access difficult +- Complicates authorization logic + +### Solution +``` +✅ Good Examples: +/tasks/678 # Direct access to task +/tasks/678/comments # Task comments +/users/012/tasks # User's tasks +/projects/345?team=789 # Project filtering +``` + +## 5. Inconsistent Error Response Formats + +### Anti-Pattern +Different error response structures across endpoints. + +```json +❌ Bad Examples: +# Endpoint 1 +{"error": "Invalid email"} + +# Endpoint 2 +{"success": false, "msg": "User not found", "code": 404} + +# Endpoint 3 +{"errors": [{"field": "name", "message": "Required"}]} +``` + +### Why It's Bad +- Makes error handling complex for clients +- Reduces code reusability +- Poor developer experience + +### Solution +```json +✅ Standardized Error Format: +{ + "error": { + "code": "VALIDATION_ERROR", + "message": "The request contains invalid data", + "details": [ + { + "field": "email", + "code": "INVALID_FORMAT", + "message": "Email address is not valid" + } + ], + "requestId": "req-123456", + "timestamp": "2024-02-16T13:00:00Z" + } +} +``` + +## 6. Missing or Poor Pagination + +### Anti-Pattern +Returning all results in a single response or inconsistent pagination. + +```json +❌ Bad Examples: +# No pagination (returns 10,000 records) +GET /api/users + +# Inconsistent pagination parameters +GET /api/users?page=1&size=10 +GET /api/orders?offset=0&limit=20 +GET /api/products?start=0&count=50 +``` + +### Why It's Bad +- Can cause performance issues +- May overwhelm clients +- Inconsistent pagination parameters confuse developers +- No way to estimate total results + +### Solution +```json +✅ Good Example: +GET /api/users?page=1&pageSize=10 + +{ + "data": [...], + "pagination": { + "page": 1, + "pageSize": 10, + "total": 150, + "totalPages": 15, + "hasNext": true, + "hasPrev": false + } +} +``` + +## 7. Exposing Internal Implementation Details + +### Anti-Pattern +URLs and field names that reflect database structure or internal architecture. + +``` +❌ Bad Examples: +/api/user_table/123 +/api/db_orders +/api/legacy_customer_data +/api/temp_migration_users + +Response fields: +{ + "user_id_pk": 123, + "internal_ref_code": "usr_abc", + "db_created_timestamp": 1645123456 +} +``` + +### Why It's Bad +- Couples API to internal implementation +- Makes refactoring difficult +- Exposes unnecessary technical details +- Reduces API longevity + +### Solution +``` +✅ Good Examples: +/api/users/123 +/api/orders +/api/customers + +Response fields: +{ + "id": 123, + "referenceCode": "usr_abc", + "createdAt": "2024-02-16T13:00:00Z" +} +``` + +## 8. Overloading Single Endpoint + +### Anti-Pattern +Using one endpoint for multiple unrelated operations based on request parameters. + +``` +❌ Bad Example: +POST /api/user-actions +{ + "action": "create_user", + "userData": {...} +} + +POST /api/user-actions +{ + "action": "delete_user", + "userId": 123 +} + +POST /api/user-actions +{ + "action": "send_email", + "userId": 123, + "emailType": "welcome" +} +``` + +### Why It's Bad +- Breaks REST principles +- Makes documentation complex +- Complicates client implementation +- Reduces discoverability + +### Solution +``` +✅ Good Examples: +POST /api/users # Create user +DELETE /api/users/123 # Delete user +POST /api/users/123/emails # Send email to user +``` + +## 9. Lack of Versioning Strategy + +### Anti-Pattern +Making breaking changes without version management. + +``` +❌ Bad Examples: +# Original API +{ + "name": "John Doe", + "age": 30 +} + +# Later (breaking change with no versioning) +{ + "firstName": "John", + "lastName": "Doe", + "birthDate": "1994-02-16" +} +``` + +### Why It's Bad +- Breaks existing clients +- Forces all clients to update simultaneously +- No graceful migration path +- Reduces API stability + +### Solution +``` +✅ Good Examples: +# Version 1 +GET /api/v1/users/123 +{ + "name": "John Doe", + "age": 30 +} + +# Version 2 (with both versions supported) +GET /api/v2/users/123 +{ + "firstName": "John", + "lastName": "Doe", + "birthDate": "1994-02-16", + "age": 30 // Backwards compatibility +} +``` + +## 10. Poor Error Messages + +### Anti-Pattern +Vague, unhelpful, or technical error messages. + +```json +❌ Bad Examples: +{"error": "Something went wrong"} +{"error": "Invalid input"} +{"error": "SQL constraint violation: FK_user_profile_id"} +{"error": "NullPointerException at line 247"} +``` + +### Why It's Bad +- Doesn't help developers fix issues +- Increases support burden +- Poor developer experience +- May expose sensitive information + +### Solution +```json +✅ Good Examples: +{ + "error": { + "code": "VALIDATION_ERROR", + "message": "The email address is required and must be in a valid format", + "details": [ + { + "field": "email", + "code": "REQUIRED", + "message": "Email address is required" + } + ] + } +} +``` + +## 11. Ignoring Content Negotiation + +### Anti-Pattern +Hard-coding response format without considering client preferences. + +``` +❌ Bad Example: +# Always returns JSON regardless of Accept header +GET /api/users/123 +Accept: application/xml +# Returns JSON anyway +``` + +### Why It's Bad +- Reduces API flexibility +- Ignores HTTP standards +- Makes integration harder for diverse clients + +### Solution +``` +✅ Good Example: +GET /api/users/123 +Accept: application/xml + +HTTP/1.1 200 OK +Content-Type: application/xml + + + + 123 + John Doe + +``` + +## 12. Stateful API Design + +### Anti-Pattern +Maintaining session state on the server between requests. + +``` +❌ Bad Example: +# Step 1: Initialize session +POST /api/session/init + +# Step 2: Set context (requires step 1) +POST /api/session/set-user/123 + +# Step 3: Get data (requires steps 1 & 2) +GET /api/session/user-data +``` + +### Why It's Bad +- Breaks REST statelessness principle +- Reduces scalability +- Makes caching difficult +- Complicates error recovery + +### Solution +``` +✅ Good Example: +# Self-contained requests +GET /api/users/123/data +Authorization: Bearer jwt-token-with-context +``` + +## 13. Inconsistent HTTP Method Usage + +### Anti-Pattern +Using HTTP methods inappropriately or inconsistently. + +``` +❌ Bad Examples: +GET /api/users/123/delete # DELETE operation with GET +POST /api/users/123/get # GET operation with POST +PUT /api/users # Creating with PUT on collection +GET /api/users/search # Search with side effects +``` + +### Why It's Bad +- Violates HTTP semantics +- Breaks caching and idempotency expectations +- Confuses developers and tools + +### Solution +``` +✅ Good Examples: +DELETE /api/users/123 # Delete with DELETE +GET /api/users/123 # Get with GET +POST /api/users # Create on collection +GET /api/users?q=search # Safe search with GET +``` + +## 14. Missing Rate Limiting Information + +### Anti-Pattern +Not providing rate limiting information to clients. + +``` +❌ Bad Example: +HTTP/1.1 429 Too Many Requests +{ + "error": "Rate limit exceeded" +} +``` + +### Why It's Bad +- Clients don't know when to retry +- No information about current limits +- Difficult to implement proper backoff strategies + +### Solution +``` +✅ Good Example: +HTTP/1.1 429 Too Many Requests +X-RateLimit-Limit: 1000 +X-RateLimit-Remaining: 0 +X-RateLimit-Reset: 1640995200 +Retry-After: 3600 + +{ + "error": { + "code": "RATE_LIMIT_EXCEEDED", + "message": "API rate limit exceeded", + "retryAfter": 3600 + } +} +``` + +## 15. Chatty API Design + +### Anti-Pattern +Requiring multiple API calls to accomplish common tasks. + +``` +❌ Bad Example: +# Get user profile requires 4 API calls +GET /api/users/123 # Basic info +GET /api/users/123/profile # Profile details +GET /api/users/123/settings # User settings +GET /api/users/123/stats # User statistics +``` + +### Why It's Bad +- Increases latency +- Creates network overhead +- Makes mobile apps inefficient +- Complicates client implementation + +### Solution +``` +✅ Good Examples: +# Single call with expansion +GET /api/users/123?include=profile,settings,stats + +# Or provide composite endpoints +GET /api/users/123/dashboard + +# Or batch operations +POST /api/batch +{ + "requests": [ + {"method": "GET", "url": "/users/123"}, + {"method": "GET", "url": "/users/123/profile"} + ] +} +``` + +## 16. No Input Validation + +### Anti-Pattern +Accepting and processing invalid input without proper validation. + +```json +❌ Bad Example: +POST /api/users +{ + "email": "not-an-email", + "age": -5, + "name": "" +} + +# API processes this and fails later or stores invalid data +``` + +### Why It's Bad +- Leads to data corruption +- Security vulnerabilities +- Difficult to debug issues +- Poor user experience + +### Solution +```json +✅ Good Example: +POST /api/users +{ + "email": "not-an-email", + "age": -5, + "name": "" +} + +HTTP/1.1 400 Bad Request +{ + "error": { + "code": "VALIDATION_ERROR", + "message": "The request contains invalid data", + "details": [ + { + "field": "email", + "code": "INVALID_FORMAT", + "message": "Email must be a valid email address" + }, + { + "field": "age", + "code": "INVALID_RANGE", + "message": "Age must be between 0 and 150" + }, + { + "field": "name", + "code": "REQUIRED", + "message": "Name is required and cannot be empty" + } + ] + } +} +``` + +## 17. Synchronous Long-Running Operations + +### Anti-Pattern +Blocking the client with long-running operations in synchronous endpoints. + +``` +❌ Bad Example: +POST /api/reports/generate +# Client waits 30 seconds for response +``` + +### Why It's Bad +- Poor user experience +- Timeouts and connection issues +- Resource waste on client and server +- Doesn't scale well + +### Solution +``` +✅ Good Example: +# Async pattern +POST /api/reports +HTTP/1.1 202 Accepted +Location: /api/reports/job-123 +{ + "jobId": "job-123", + "status": "processing", + "estimatedCompletion": "2024-02-16T13:05:00Z" +} + +# Check status +GET /api/reports/job-123 +{ + "jobId": "job-123", + "status": "completed", + "result": "/api/reports/download/report-456" +} +``` + +## Prevention Strategies + +### 1. API Design Reviews +- Implement mandatory design reviews +- Use checklists based on these anti-patterns +- Include multiple stakeholders + +### 2. API Style Guides +- Create and enforce API style guides +- Use linting tools for consistency +- Regular training for development teams + +### 3. Automated Testing +- Test for common anti-patterns +- Include contract testing +- Monitor API usage patterns + +### 4. Documentation Standards +- Require comprehensive API documentation +- Include examples and error scenarios +- Keep documentation up-to-date + +### 5. Client Feedback +- Regularly collect feedback from API consumers +- Monitor API usage analytics +- Conduct developer experience surveys + +## Conclusion + +Avoiding these anti-patterns requires: +- Understanding REST principles +- Consistent design standards +- Regular review and refactoring +- Focus on developer experience +- Proper tooling and automation + +Remember: A well-designed API is an asset that grows in value over time, while a poorly designed API becomes a liability that hampers development and adoption. \ No newline at end of file diff --git a/skills/api-design-reviewer/references/rest_design_rules.md b/skills/api-design-reviewer/references/rest_design_rules.md new file mode 100644 index 00000000..1eb9b1f0 --- /dev/null +++ b/skills/api-design-reviewer/references/rest_design_rules.md @@ -0,0 +1,487 @@ +# REST API Design Rules Reference + +## Core Principles + +### 1. Resources, Not Actions +REST APIs should focus on **resources** (nouns) rather than **actions** (verbs). The HTTP methods provide the actions. + +``` +✅ Good: +GET /users # Get all users +GET /users/123 # Get user 123 +POST /users # Create new user +PUT /users/123 # Update user 123 +DELETE /users/123 # Delete user 123 + +❌ Bad: +POST /getUsers +POST /createUser +POST /updateUser/123 +POST /deleteUser/123 +``` + +### 2. Hierarchical Resource Structure +Use hierarchical URLs to represent resource relationships: + +``` +/users/123/orders/456/items/789 +``` + +But avoid excessive nesting (max 3-4 levels): + +``` +❌ Too deep: /companies/123/departments/456/teams/789/members/012/tasks/345 +✅ Better: /tasks/345?member=012&team=789 +``` + +## Resource Naming Conventions + +### URLs Should Use Kebab-Case +``` +✅ Good: +/user-profiles +/order-items +/shipping-addresses + +❌ Bad: +/userProfiles +/user_profiles +/orderItems +``` + +### Collections vs Individual Resources +``` +Collection: /users +Individual: /users/123 +Sub-resource: /users/123/orders +``` + +### Pluralization Rules +- Use **plural nouns** for collections: `/users`, `/orders` +- Use **singular nouns** for single resources: `/user-profile`, `/current-session` +- Be consistent throughout your API + +## HTTP Methods Usage + +### GET - Safe and Idempotent +- **Purpose**: Retrieve data +- **Safe**: No side effects +- **Idempotent**: Multiple calls return same result +- **Request Body**: Should not have one +- **Cacheable**: Yes + +``` +GET /users/123 +GET /users?status=active&limit=10 +``` + +### POST - Not Idempotent +- **Purpose**: Create resources, non-idempotent operations +- **Safe**: No +- **Idempotent**: No +- **Request Body**: Usually required +- **Cacheable**: Generally no + +``` +POST /users # Create new user +POST /users/123/activate # Activate user (action) +``` + +### PUT - Idempotent +- **Purpose**: Create or completely replace a resource +- **Safe**: No +- **Idempotent**: Yes +- **Request Body**: Required (complete resource) +- **Cacheable**: No + +``` +PUT /users/123 # Replace entire user resource +``` + +### PATCH - Partial Update +- **Purpose**: Partially update a resource +- **Safe**: No +- **Idempotent**: Not necessarily +- **Request Body**: Required (partial resource) +- **Cacheable**: No + +``` +PATCH /users/123 # Update only specified fields +``` + +### DELETE - Idempotent +- **Purpose**: Remove a resource +- **Safe**: No +- **Idempotent**: Yes (same result if called multiple times) +- **Request Body**: Usually not needed +- **Cacheable**: No + +``` +DELETE /users/123 +``` + +## Status Codes + +### Success Codes (2xx) +- **200 OK**: Standard success response +- **201 Created**: Resource created successfully (POST) +- **202 Accepted**: Request accepted for processing (async) +- **204 No Content**: Success with no response body (DELETE, PUT) + +### Redirection Codes (3xx) +- **301 Moved Permanently**: Resource permanently moved +- **302 Found**: Temporary redirect +- **304 Not Modified**: Use cached version + +### Client Error Codes (4xx) +- **400 Bad Request**: Invalid request syntax or data +- **401 Unauthorized**: Authentication required +- **403 Forbidden**: Access denied (user authenticated but not authorized) +- **404 Not Found**: Resource not found +- **405 Method Not Allowed**: HTTP method not supported +- **409 Conflict**: Resource conflict (duplicates, version mismatch) +- **422 Unprocessable Entity**: Valid syntax but semantic errors +- **429 Too Many Requests**: Rate limit exceeded + +### Server Error Codes (5xx) +- **500 Internal Server Error**: Unexpected server error +- **502 Bad Gateway**: Invalid response from upstream server +- **503 Service Unavailable**: Server temporarily unavailable +- **504 Gateway Timeout**: Upstream server timeout + +## URL Design Patterns + +### Query Parameters for Filtering +``` +GET /users?status=active +GET /users?role=admin&department=engineering +GET /orders?created_after=2024-01-01&status=pending +``` + +### Pagination Parameters +``` +# Offset-based +GET /users?offset=20&limit=10 + +# Cursor-based +GET /users?cursor=eyJpZCI6MTIzfQ&limit=10 + +# Page-based +GET /users?page=3&page_size=10 +``` + +### Sorting Parameters +``` +GET /users?sort=created_at # Ascending +GET /users?sort=-created_at # Descending (prefix with -) +GET /users?sort=last_name,first_name # Multiple fields +``` + +### Field Selection +``` +GET /users?fields=id,name,email +GET /users/123?include=orders,profile +GET /users/123?exclude=internal_notes +``` + +### Search Parameters +``` +GET /users?q=john +GET /products?search=laptop&category=electronics +``` + +## Response Format Standards + +### Consistent Response Structure +```json +{ + "data": { + "id": 123, + "name": "John Doe", + "email": "john@example.com" + }, + "meta": { + "timestamp": "2024-02-16T13:00:00Z", + "version": "1.0" + } +} +``` + +### Collection Responses +```json +{ + "data": [ + {"id": 1, "name": "Item 1"}, + {"id": 2, "name": "Item 2"} + ], + "pagination": { + "total": 150, + "page": 1, + "pageSize": 10, + "totalPages": 15, + "hasNext": true, + "hasPrev": false + }, + "meta": { + "timestamp": "2024-02-16T13:00:00Z" + } +} +``` + +### Error Response Format +```json +{ + "error": { + "code": "VALIDATION_ERROR", + "message": "The request contains invalid parameters", + "details": [ + { + "field": "email", + "code": "INVALID_FORMAT", + "message": "Email address is not valid" + } + ], + "requestId": "req-123456", + "timestamp": "2024-02-16T13:00:00Z" + } +} +``` + +## Field Naming Conventions + +### Use camelCase for JSON Fields +```json +✅ Good: +{ + "firstName": "John", + "lastName": "Doe", + "createdAt": "2024-02-16T13:00:00Z", + "isActive": true +} + +❌ Bad: +{ + "first_name": "John", + "LastName": "Doe", + "created-at": "2024-02-16T13:00:00Z" +} +``` + +### Boolean Fields +Use positive, clear names with "is", "has", "can", or "should" prefixes: + +```json +✅ Good: +{ + "isActive": true, + "hasPermission": false, + "canEdit": true, + "shouldNotify": false +} + +❌ Bad: +{ + "active": true, + "disabled": false, // Double negative + "permission": false // Unclear meaning +} +``` + +### Date/Time Fields +- Use ISO 8601 format: `2024-02-16T13:00:00Z` +- Include timezone information +- Use consistent field naming: + +```json +{ + "createdAt": "2024-02-16T13:00:00Z", + "updatedAt": "2024-02-16T13:30:00Z", + "deletedAt": null, + "publishedAt": "2024-02-16T14:00:00Z" +} +``` + +## Content Negotiation + +### Accept Headers +``` +Accept: application/json +Accept: application/xml +Accept: application/json; version=1 +``` + +### Content-Type Headers +``` +Content-Type: application/json +Content-Type: application/json; charset=utf-8 +Content-Type: multipart/form-data +``` + +### Versioning via Headers +``` +Accept: application/vnd.myapi.v1+json +API-Version: 1.0 +``` + +## Caching Guidelines + +### Cache-Control Headers +``` +Cache-Control: public, max-age=3600 # Cache for 1 hour +Cache-Control: private, max-age=0 # Don't cache +Cache-Control: no-cache, must-revalidate # Always validate +``` + +### ETags for Conditional Requests +``` +HTTP/1.1 200 OK +ETag: "123456789" +Last-Modified: Wed, 21 Oct 2015 07:28:00 GMT + +# Client subsequent request: +If-None-Match: "123456789" +If-Modified-Since: Wed, 21 Oct 2015 07:28:00 GMT +``` + +## Security Headers + +### Authentication +``` +Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9... +Authorization: Basic dXNlcjpwYXNzd29yZA== +Authorization: Api-Key abc123def456 +``` + +### CORS Headers +``` +Access-Control-Allow-Origin: https://example.com +Access-Control-Allow-Methods: GET, POST, PUT, DELETE +Access-Control-Allow-Headers: Content-Type, Authorization +``` + +## Rate Limiting + +### Rate Limit Headers +``` +X-RateLimit-Limit: 1000 +X-RateLimit-Remaining: 999 +X-RateLimit-Reset: 1640995200 +X-RateLimit-Window: 3600 +``` + +### Rate Limit Exceeded Response +```json +HTTP/1.1 429 Too Many Requests +Retry-After: 3600 + +{ + "error": { + "code": "RATE_LIMIT_EXCEEDED", + "message": "API rate limit exceeded", + "details": { + "limit": 1000, + "window": "1 hour", + "retryAfter": 3600 + } + } +} +``` + +## Hypermedia (HATEOAS) + +### Links in Responses +```json +{ + "id": 123, + "name": "John Doe", + "email": "john@example.com", + "_links": { + "self": { + "href": "/users/123" + }, + "orders": { + "href": "/users/123/orders" + }, + "edit": { + "href": "/users/123", + "method": "PUT" + }, + "delete": { + "href": "/users/123", + "method": "DELETE" + } + } +} +``` + +### Link Relations +- **self**: Link to the resource itself +- **edit**: Link to edit the resource +- **delete**: Link to delete the resource +- **related**: Link to related resources +- **next/prev**: Pagination links + +## Common Anti-Patterns to Avoid + +### 1. Verbs in URLs +``` +❌ Bad: /api/getUser/123 +✅ Good: GET /api/users/123 +``` + +### 2. Inconsistent Naming +``` +❌ Bad: /user-profiles and /userAddresses +✅ Good: /user-profiles and /user-addresses +``` + +### 3. Deep Nesting +``` +❌ Bad: /companies/123/departments/456/teams/789/members/012 +✅ Good: /team-members/012?team=789 +``` + +### 4. Ignoring HTTP Status Codes +``` +❌ Bad: Always return 200 with error info in body +✅ Good: Use appropriate status codes (404, 400, 500, etc.) +``` + +### 5. Exposing Internal Structure +``` +❌ Bad: /api/database_table_users +✅ Good: /api/users +``` + +### 6. No Versioning Strategy +``` +❌ Bad: Breaking changes without version management +✅ Good: /api/v1/users or Accept: application/vnd.api+json;version=1 +``` + +### 7. Inconsistent Error Responses +``` +❌ Bad: Different error formats for different endpoints +✅ Good: Standardized error response structure +``` + +## Best Practices Summary + +1. **Use nouns for resources, not verbs** +2. **Leverage HTTP methods correctly** +3. **Maintain consistent naming conventions** +4. **Implement proper error handling** +5. **Use appropriate HTTP status codes** +6. **Design for cacheability** +7. **Implement security from the start** +8. **Plan for versioning** +9. **Provide comprehensive documentation** +10. **Follow HATEOAS principles when applicable** + +## Further Reading + +- [RFC 7231 - HTTP/1.1 Semantics and Content](https://tools.ietf.org/html/rfc7231) +- [RFC 6570 - URI Template](https://tools.ietf.org/html/rfc6570) +- [OpenAPI Specification](https://swagger.io/specification/) +- [REST API Design Best Practices](https://www.restapitutorial.com/) +- [HTTP Status Code Definitions](https://httpstatuses.com/) \ No newline at end of file diff --git a/skills/api-design-reviewer/scripts/api_linter.py b/skills/api-design-reviewer/scripts/api_linter.py new file mode 100644 index 00000000..53637d5e --- /dev/null +++ b/skills/api-design-reviewer/scripts/api_linter.py @@ -0,0 +1,914 @@ +#!/usr/bin/env python3 +""" +API Linter - Analyzes OpenAPI/Swagger specifications for REST conventions and best practices. + +This script validates API designs against established conventions including: +- Resource naming conventions (kebab-case resources, camelCase fields) +- HTTP method usage patterns +- URL structure consistency +- Error response format standards +- Documentation completeness +- Pagination patterns +- Versioning compliance + +Supports both OpenAPI JSON specifications and raw endpoint definition JSON. +""" + +import argparse +import json +import re +import sys +from typing import Any, Dict, List, Tuple, Optional, Set +from urllib.parse import urlparse +from dataclasses import dataclass, field + + +@dataclass +class LintIssue: + """Represents a linting issue found in the API specification.""" + severity: str # 'error', 'warning', 'info' + category: str + message: str + path: str + suggestion: str = "" + line_number: Optional[int] = None + + +@dataclass +class LintReport: + """Complete linting report with issues and statistics.""" + issues: List[LintIssue] = field(default_factory=list) + total_endpoints: int = 0 + endpoints_with_issues: int = 0 + score: float = 0.0 + + def add_issue(self, issue: LintIssue) -> None: + """Add an issue to the report.""" + self.issues.append(issue) + + def get_issues_by_severity(self) -> Dict[str, List[LintIssue]]: + """Group issues by severity level.""" + grouped = {'error': [], 'warning': [], 'info': []} + for issue in self.issues: + if issue.severity in grouped: + grouped[issue.severity].append(issue) + return grouped + + def calculate_score(self) -> float: + """Calculate overall API quality score (0-100).""" + if self.total_endpoints == 0: + return 100.0 + + error_penalty = len([i for i in self.issues if i.severity == 'error']) * 10 + warning_penalty = len([i for i in self.issues if i.severity == 'warning']) * 3 + info_penalty = len([i for i in self.issues if i.severity == 'info']) * 1 + + total_penalty = error_penalty + warning_penalty + info_penalty + base_score = 100.0 + + # Penalty per endpoint to normalize across API sizes + penalty_per_endpoint = total_penalty / self.total_endpoints if self.total_endpoints > 0 else total_penalty + + self.score = max(0.0, base_score - penalty_per_endpoint) + return self.score + + +class APILinter: + """Main API linting engine.""" + + def __init__(self): + self.report = LintReport() + self.openapi_spec: Optional[Dict] = None + self.raw_endpoints: Optional[Dict] = None + + # Regex patterns for naming conventions + self.kebab_case_pattern = re.compile(r'^[a-z]+(?:-[a-z0-9]+)*$') + self.camel_case_pattern = re.compile(r'^[a-z][a-zA-Z0-9]*$') + self.snake_case_pattern = re.compile(r'^[a-z]+(?:_[a-z0-9]+)*$') + self.pascal_case_pattern = re.compile(r'^[A-Z][a-zA-Z0-9]*$') + + # Standard HTTP methods + self.http_methods = {'GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'HEAD', 'OPTIONS'} + + # Standard HTTP status codes by method + self.standard_status_codes = { + 'GET': {200, 304, 404}, + 'POST': {200, 201, 400, 409, 422}, + 'PUT': {200, 204, 400, 404, 409}, + 'PATCH': {200, 204, 400, 404, 409}, + 'DELETE': {200, 204, 404}, + 'HEAD': {200, 404}, + 'OPTIONS': {200} + } + + # Common error status codes + self.common_error_codes = {400, 401, 403, 404, 405, 409, 422, 429, 500, 502, 503} + + def lint_openapi_spec(self, spec: Dict[str, Any]) -> LintReport: + """Lint an OpenAPI/Swagger specification.""" + self.openapi_spec = spec + self.report = LintReport() + + # Basic structure validation + self._validate_openapi_structure() + + # Info section validation + self._validate_info_section() + + # Server section validation + self._validate_servers_section() + + # Paths validation (main linting logic) + self._validate_paths_section() + + # Components validation + self._validate_components_section() + + # Security validation + self._validate_security_section() + + # Calculate final score + self.report.calculate_score() + + return self.report + + def lint_raw_endpoints(self, endpoints: Dict[str, Any]) -> LintReport: + """Lint raw endpoint definitions.""" + self.raw_endpoints = endpoints + self.report = LintReport() + + # Validate raw endpoint structure + self._validate_raw_endpoint_structure() + + # Lint each endpoint + for endpoint_path, endpoint_data in endpoints.get('endpoints', {}).items(): + self._lint_raw_endpoint(endpoint_path, endpoint_data) + + self.report.calculate_score() + return self.report + + def _validate_openapi_structure(self) -> None: + """Validate basic OpenAPI document structure.""" + required_fields = ['openapi', 'info', 'paths'] + + for field in required_fields: + if field not in self.openapi_spec: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message=f"Missing required field: {field}", + path=f"/{field}", + suggestion=f"Add the '{field}' field to the root of your OpenAPI specification" + )) + + def _validate_info_section(self) -> None: + """Validate the info section of OpenAPI spec.""" + if 'info' not in self.openapi_spec: + return + + info = self.openapi_spec['info'] + required_info_fields = ['title', 'version'] + recommended_info_fields = ['description', 'contact'] + + for field in required_info_fields: + if field not in info: + self.report.add_issue(LintIssue( + severity='error', + category='documentation', + message=f"Missing required info field: {field}", + path=f"/info/{field}", + suggestion=f"Add a '{field}' field to the info section" + )) + + for field in recommended_info_fields: + if field not in info: + self.report.add_issue(LintIssue( + severity='warning', + category='documentation', + message=f"Missing recommended info field: {field}", + path=f"/info/{field}", + suggestion=f"Consider adding a '{field}' field to improve API documentation" + )) + + # Validate version format + if 'version' in info: + version = info['version'] + if not re.match(r'^\d+\.\d+(\.\d+)?(-\w+)?$', version): + self.report.add_issue(LintIssue( + severity='warning', + category='versioning', + message=f"Version format '{version}' doesn't follow semantic versioning", + path="/info/version", + suggestion="Use semantic versioning format (e.g., '1.0.0', '2.1.3-beta')" + )) + + def _validate_servers_section(self) -> None: + """Validate the servers section.""" + if 'servers' not in self.openapi_spec: + self.report.add_issue(LintIssue( + severity='warning', + category='configuration', + message="Missing servers section", + path="/servers", + suggestion="Add a servers section to specify API base URLs" + )) + return + + servers = self.openapi_spec['servers'] + if not isinstance(servers, list) or len(servers) == 0: + self.report.add_issue(LintIssue( + severity='warning', + category='configuration', + message="Empty servers section", + path="/servers", + suggestion="Add at least one server URL" + )) + + def _validate_paths_section(self) -> None: + """Validate all API paths and operations.""" + if 'paths' not in self.openapi_spec: + return + + paths = self.openapi_spec['paths'] + if not paths: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message="No paths defined in API specification", + path="/paths", + suggestion="Define at least one API endpoint" + )) + return + + self.report.total_endpoints = sum( + len([method for method in path_obj.keys() if method.upper() in self.http_methods]) + for path_obj in paths.values() if isinstance(path_obj, dict) + ) + + endpoints_with_issues = set() + + for path, path_obj in paths.items(): + if not isinstance(path_obj, dict): + continue + + # Validate path structure + path_issues = self._validate_path_structure(path) + if path_issues: + endpoints_with_issues.add(path) + + # Validate each operation in the path + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + operation_issues = self._validate_operation(path, method.upper(), operation) + if operation_issues: + endpoints_with_issues.add(path) + + self.report.endpoints_with_issues = len(endpoints_with_issues) + + def _validate_path_structure(self, path: str) -> bool: + """Validate REST path structure and naming conventions.""" + has_issues = False + + # Check if path starts with slash + if not path.startswith('/'): + self.report.add_issue(LintIssue( + severity='error', + category='url_structure', + message=f"Path must start with '/' character: {path}", + path=f"/paths/{path}", + suggestion=f"Change '{path}' to '/{path.lstrip('/')}'" + )) + has_issues = True + + # Split path into segments + segments = [seg for seg in path.split('/') if seg] + + # Check for empty segments (double slashes) + if '//' in path: + self.report.add_issue(LintIssue( + severity='error', + category='url_structure', + message=f"Path contains empty segments: {path}", + path=f"/paths/{path}", + suggestion="Remove double slashes from the path" + )) + has_issues = True + + # Validate each segment + for i, segment in enumerate(segments): + # Skip parameter segments + if segment.startswith('{') and segment.endswith('}'): + # Validate parameter naming + param_name = segment[1:-1] + if not self.camel_case_pattern.match(param_name) and not self.kebab_case_pattern.match(param_name): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Path parameter '{param_name}' should use camelCase or kebab-case", + path=f"/paths/{path}", + suggestion=f"Use camelCase (e.g., 'userId') or kebab-case (e.g., 'user-id')" + )) + has_issues = True + continue + + # Check for resource naming conventions + if not self.kebab_case_pattern.match(segment): + # Allow version segments like 'v1', 'v2' + if not re.match(r'^v\d+$', segment): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Resource segment '{segment}' should use kebab-case", + path=f"/paths/{path}", + suggestion=f"Use kebab-case for '{segment}' (e.g., 'user-profiles', 'order-items')" + )) + has_issues = True + + # Check for verb usage in URLs (anti-pattern) + common_verbs = {'get', 'post', 'put', 'delete', 'create', 'update', 'remove', 'add'} + if segment.lower() in common_verbs: + self.report.add_issue(LintIssue( + severity='warning', + category='rest_conventions', + message=f"Avoid verbs in URLs: '{segment}' in {path}", + path=f"/paths/{path}", + suggestion="Use HTTP methods instead of verbs in URLs. Use nouns for resources." + )) + has_issues = True + + # Check path depth (avoid over-nesting) + if len(segments) > 6: + self.report.add_issue(LintIssue( + severity='warning', + category='url_structure', + message=f"Path has excessive nesting ({len(segments)} levels): {path}", + path=f"/paths/{path}", + suggestion="Consider flattening the resource hierarchy or using query parameters" + )) + has_issues = True + + # Check for consistent versioning + if any('v' + str(i) in segments for i in range(1, 10)): + version_segments = [seg for seg in segments if re.match(r'^v\d+$', seg)] + if len(version_segments) > 1: + self.report.add_issue(LintIssue( + severity='error', + category='versioning', + message=f"Multiple version segments in path: {path}", + path=f"/paths/{path}", + suggestion="Use only one version segment per path" + )) + has_issues = True + + return has_issues + + def _validate_operation(self, path: str, method: str, operation: Dict[str, Any]) -> bool: + """Validate individual operation (HTTP method + path combination).""" + has_issues = False + operation_path = f"/paths/{path}/{method.lower()}" + + # Check for required operation fields + if 'responses' not in operation: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message=f"Missing responses section for {method} {path}", + path=f"{operation_path}/responses", + suggestion="Define expected responses for this operation" + )) + has_issues = True + + # Check for operation documentation + if 'summary' not in operation: + self.report.add_issue(LintIssue( + severity='warning', + category='documentation', + message=f"Missing summary for {method} {path}", + path=f"{operation_path}/summary", + suggestion="Add a brief summary describing what this operation does" + )) + has_issues = True + + if 'description' not in operation: + self.report.add_issue(LintIssue( + severity='info', + category='documentation', + message=f"Missing description for {method} {path}", + path=f"{operation_path}/description", + suggestion="Add a detailed description for better API documentation" + )) + has_issues = True + + # Validate HTTP method usage patterns + method_issues = self._validate_http_method_usage(path, method, operation) + if method_issues: + has_issues = True + + # Validate responses + if 'responses' in operation: + response_issues = self._validate_responses(path, method, operation['responses']) + if response_issues: + has_issues = True + + # Validate parameters + if 'parameters' in operation: + param_issues = self._validate_parameters(path, method, operation['parameters']) + if param_issues: + has_issues = True + + # Validate request body + if 'requestBody' in operation: + body_issues = self._validate_request_body(path, method, operation['requestBody']) + if body_issues: + has_issues = True + + return has_issues + + def _validate_http_method_usage(self, path: str, method: str, operation: Dict[str, Any]) -> bool: + """Validate proper HTTP method usage patterns.""" + has_issues = False + + # GET requests should not have request body + if method == 'GET' and 'requestBody' in operation: + self.report.add_issue(LintIssue( + severity='error', + category='rest_conventions', + message=f"GET request should not have request body: {method} {path}", + path=f"/paths/{path}/{method.lower()}/requestBody", + suggestion="Remove requestBody from GET request or use POST if body is needed" + )) + has_issues = True + + # DELETE requests typically should not have request body + if method == 'DELETE' and 'requestBody' in operation: + self.report.add_issue(LintIssue( + severity='warning', + category='rest_conventions', + message=f"DELETE request typically should not have request body: {method} {path}", + path=f"/paths/{path}/{method.lower()}/requestBody", + suggestion="Consider using query parameters or path parameters instead" + )) + has_issues = True + + # POST/PUT/PATCH should typically have request body (except for actions) + if method in ['POST', 'PUT', 'PATCH'] and 'requestBody' not in operation: + # Check if this is an action endpoint + if not any(action in path.lower() for action in ['activate', 'deactivate', 'reset', 'confirm']): + self.report.add_issue(LintIssue( + severity='info', + category='rest_conventions', + message=f"{method} request typically should have request body: {method} {path}", + path=f"/paths/{path}/{method.lower()}", + suggestion=f"Consider adding requestBody for {method} operation or use GET if no data is being sent" + )) + has_issues = True + + return has_issues + + def _validate_responses(self, path: str, method: str, responses: Dict[str, Any]) -> bool: + """Validate response definitions.""" + has_issues = False + + # Check for success response + success_codes = {'200', '201', '202', '204'} + has_success = any(code in responses for code in success_codes) + + if not has_success: + self.report.add_issue(LintIssue( + severity='error', + category='responses', + message=f"Missing success response for {method} {path}", + path=f"/paths/{path}/{method.lower()}/responses", + suggestion="Define at least one success response (200, 201, 202, or 204)" + )) + has_issues = True + + # Check for error responses + has_error_responses = any(code.startswith('4') or code.startswith('5') for code in responses.keys()) + + if not has_error_responses: + self.report.add_issue(LintIssue( + severity='warning', + category='responses', + message=f"Missing error responses for {method} {path}", + path=f"/paths/{path}/{method.lower()}/responses", + suggestion="Define common error responses (400, 404, 500, etc.)" + )) + has_issues = True + + # Validate individual response codes + for status_code, response in responses.items(): + if status_code == 'default': + continue + + try: + code_int = int(status_code) + except ValueError: + self.report.add_issue(LintIssue( + severity='error', + category='responses', + message=f"Invalid status code '{status_code}' for {method} {path}", + path=f"/paths/{path}/{method.lower()}/responses/{status_code}", + suggestion="Use valid HTTP status codes (e.g., 200, 404, 500)" + )) + has_issues = True + continue + + # Check if status code is appropriate for the method + expected_codes = self.standard_status_codes.get(method, set()) + common_codes = {400, 401, 403, 404, 429, 500} # Always acceptable + + if expected_codes and code_int not in expected_codes and code_int not in common_codes: + self.report.add_issue(LintIssue( + severity='info', + category='responses', + message=f"Uncommon status code {status_code} for {method} {path}", + path=f"/paths/{path}/{method.lower()}/responses/{status_code}", + suggestion=f"Consider using standard codes for {method}: {sorted(expected_codes)}" + )) + has_issues = True + + return has_issues + + def _validate_parameters(self, path: str, method: str, parameters: List[Dict[str, Any]]) -> bool: + """Validate parameter definitions.""" + has_issues = False + + for i, param in enumerate(parameters): + param_path = f"/paths/{path}/{method.lower()}/parameters[{i}]" + + # Check required fields + if 'name' not in param: + self.report.add_issue(LintIssue( + severity='error', + category='parameters', + message=f"Parameter missing name field in {method} {path}", + path=f"{param_path}/name", + suggestion="Add a name field to the parameter" + )) + has_issues = True + continue + + if 'in' not in param: + self.report.add_issue(LintIssue( + severity='error', + category='parameters', + message=f"Parameter '{param['name']}' missing 'in' field in {method} {path}", + path=f"{param_path}/in", + suggestion="Specify parameter location (query, path, header, cookie)" + )) + has_issues = True + + # Validate parameter naming + param_name = param['name'] + param_location = param.get('in', '') + + if param_location == 'query': + # Query parameters should use camelCase or kebab-case + if not self.camel_case_pattern.match(param_name) and not self.kebab_case_pattern.match(param_name): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Query parameter '{param_name}' should use camelCase or kebab-case in {method} {path}", + path=f"{param_path}/name", + suggestion="Use camelCase (e.g., 'pageSize') or kebab-case (e.g., 'page-size')" + )) + has_issues = True + + elif param_location == 'path': + # Path parameters should use camelCase or kebab-case + if not self.camel_case_pattern.match(param_name) and not self.kebab_case_pattern.match(param_name): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Path parameter '{param_name}' should use camelCase or kebab-case in {method} {path}", + path=f"{param_path}/name", + suggestion="Use camelCase (e.g., 'userId') or kebab-case (e.g., 'user-id')" + )) + has_issues = True + + # Path parameters must be required + if not param.get('required', False): + self.report.add_issue(LintIssue( + severity='error', + category='parameters', + message=f"Path parameter '{param_name}' must be required in {method} {path}", + path=f"{param_path}/required", + suggestion="Set required: true for path parameters" + )) + has_issues = True + + return has_issues + + def _validate_request_body(self, path: str, method: str, request_body: Dict[str, Any]) -> bool: + """Validate request body definition.""" + has_issues = False + + if 'content' not in request_body: + self.report.add_issue(LintIssue( + severity='error', + category='request_body', + message=f"Request body missing content for {method} {path}", + path=f"/paths/{path}/{method.lower()}/requestBody/content", + suggestion="Define content types for the request body" + )) + has_issues = True + + return has_issues + + def _validate_components_section(self) -> None: + """Validate the components section.""" + if 'components' not in self.openapi_spec: + self.report.add_issue(LintIssue( + severity='info', + category='structure', + message="Missing components section", + path="/components", + suggestion="Consider defining reusable components (schemas, responses, parameters)" + )) + return + + components = self.openapi_spec['components'] + + # Validate schemas + if 'schemas' in components: + self._validate_schemas(components['schemas']) + + def _validate_schemas(self, schemas: Dict[str, Any]) -> None: + """Validate schema definitions.""" + for schema_name, schema in schemas.items(): + # Check schema naming (should be PascalCase) + if not self.pascal_case_pattern.match(schema_name): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Schema name '{schema_name}' should use PascalCase", + path=f"/components/schemas/{schema_name}", + suggestion=f"Use PascalCase for schema names (e.g., 'UserProfile', 'OrderItem')" + )) + + # Validate schema properties + if isinstance(schema, dict) and 'properties' in schema: + self._validate_schema_properties(schema_name, schema['properties']) + + def _validate_schema_properties(self, schema_name: str, properties: Dict[str, Any]) -> None: + """Validate schema property naming.""" + for prop_name, prop_def in properties.items(): + # Properties should use camelCase + if not self.camel_case_pattern.match(prop_name): + self.report.add_issue(LintIssue( + severity='warning', + category='naming', + message=f"Property '{prop_name}' in schema '{schema_name}' should use camelCase", + path=f"/components/schemas/{schema_name}/properties/{prop_name}", + suggestion="Use camelCase for property names (e.g., 'firstName', 'createdAt')" + )) + + def _validate_security_section(self) -> None: + """Validate security definitions.""" + if 'security' not in self.openapi_spec and 'components' not in self.openapi_spec: + self.report.add_issue(LintIssue( + severity='warning', + category='security', + message="No security configuration found", + path="/security", + suggestion="Define security schemes and apply them to operations" + )) + + def _validate_raw_endpoint_structure(self) -> None: + """Validate structure of raw endpoint definitions.""" + if 'endpoints' not in self.raw_endpoints: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message="Missing 'endpoints' field in raw endpoint definition", + path="/endpoints", + suggestion="Provide an 'endpoints' object containing endpoint definitions" + )) + return + + endpoints = self.raw_endpoints['endpoints'] + self.report.total_endpoints = len(endpoints) + + def _lint_raw_endpoint(self, path: str, endpoint_data: Dict[str, Any]) -> None: + """Lint individual raw endpoint definition.""" + # Validate path structure + self._validate_path_structure(path) + + # Check for required fields + if 'method' not in endpoint_data: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message=f"Missing method field for endpoint {path}", + path=f"/endpoints/{path}/method", + suggestion="Specify HTTP method (GET, POST, PUT, PATCH, DELETE)" + )) + return + + method = endpoint_data['method'].upper() + if method not in self.http_methods: + self.report.add_issue(LintIssue( + severity='error', + category='structure', + message=f"Invalid HTTP method '{method}' for endpoint {path}", + path=f"/endpoints/{path}/method", + suggestion=f"Use valid HTTP methods: {', '.join(sorted(self.http_methods))}" + )) + + def generate_json_report(self) -> str: + """Generate JSON format report.""" + issues_by_severity = self.report.get_issues_by_severity() + + report_data = { + "summary": { + "total_endpoints": self.report.total_endpoints, + "endpoints_with_issues": self.report.endpoints_with_issues, + "total_issues": len(self.report.issues), + "errors": len(issues_by_severity['error']), + "warnings": len(issues_by_severity['warning']), + "info": len(issues_by_severity['info']), + "score": round(self.report.score, 2) + }, + "issues": [] + } + + for issue in self.report.issues: + report_data["issues"].append({ + "severity": issue.severity, + "category": issue.category, + "message": issue.message, + "path": issue.path, + "suggestion": issue.suggestion + }) + + return json.dumps(report_data, indent=2) + + def generate_text_report(self) -> str: + """Generate human-readable text report.""" + issues_by_severity = self.report.get_issues_by_severity() + + report_lines = [ + "═══════════════════════════════════════════════════════════════", + " API LINTING REPORT", + "═══════════════════════════════════════════════════════════════", + "", + "SUMMARY:", + f" Total Endpoints: {self.report.total_endpoints}", + f" Endpoints with Issues: {self.report.endpoints_with_issues}", + f" Overall Score: {self.report.score:.1f}/100.0", + "", + "ISSUE BREAKDOWN:", + f" 🔴 Errors: {len(issues_by_severity['error'])}", + f" 🟡 Warnings: {len(issues_by_severity['warning'])}", + f" ℹ️ Info: {len(issues_by_severity['info'])}", + "", + ] + + if not self.report.issues: + report_lines.extend([ + "🎉 Congratulations! No issues found in your API specification.", + "" + ]) + else: + # Group issues by category + issues_by_category = {} + for issue in self.report.issues: + if issue.category not in issues_by_category: + issues_by_category[issue.category] = [] + issues_by_category[issue.category].append(issue) + + for category, issues in issues_by_category.items(): + report_lines.append(f"{'═' * 60}") + report_lines.append(f"CATEGORY: {category.upper().replace('_', ' ')}") + report_lines.append(f"{'═' * 60}") + + for issue in issues: + severity_icon = {"error": "🔴", "warning": "🟡", "info": "ℹ️"}[issue.severity] + + report_lines.extend([ + f"{severity_icon} {issue.severity.upper()}: {issue.message}", + f" Path: {issue.path}", + ]) + + if issue.suggestion: + report_lines.append(f" 💡 Suggestion: {issue.suggestion}") + + report_lines.append("") + + # Add scoring breakdown + report_lines.extend([ + "═══════════════════════════════════════════════════════════════", + "SCORING DETAILS:", + "═══════════════════════════════════════════════════════════════", + f"Base Score: 100.0", + f"Errors Penalty: -{len(issues_by_severity['error']) * 10} (10 points per error)", + f"Warnings Penalty: -{len(issues_by_severity['warning']) * 3} (3 points per warning)", + f"Info Penalty: -{len(issues_by_severity['info']) * 1} (1 point per info)", + f"Final Score: {self.report.score:.1f}/100.0", + "" + ]) + + # Add recommendations based on score + if self.report.score >= 90: + report_lines.append("🏆 Excellent! Your API design follows best practices.") + elif self.report.score >= 80: + report_lines.append("✅ Good API design with minor areas for improvement.") + elif self.report.score >= 70: + report_lines.append("⚠️ Fair API design. Consider addressing warnings and errors.") + elif self.report.score >= 50: + report_lines.append("❌ Poor API design. Multiple issues need attention.") + else: + report_lines.append("🚨 Critical API design issues. Immediate attention required.") + + return "\n".join(report_lines) + + +def main(): + """Main CLI entry point.""" + parser = argparse.ArgumentParser( + description="Analyze OpenAPI/Swagger specifications for REST conventions and best practices", + formatter_class=argparse.RawDescriptionHelpFormatter, + epilog=""" +Examples: + python api_linter.py openapi.json + python api_linter.py --format json openapi.json > report.json + python api_linter.py --raw-endpoints endpoints.json + """ + ) + + parser.add_argument( + 'input_file', + help='Input file: OpenAPI/Swagger JSON file or raw endpoints JSON' + ) + + parser.add_argument( + '--format', + choices=['text', 'json'], + default='text', + help='Output format (default: text)' + ) + + parser.add_argument( + '--raw-endpoints', + action='store_true', + help='Treat input as raw endpoint definitions instead of OpenAPI spec' + ) + + parser.add_argument( + '--output', + help='Output file (default: stdout)' + ) + + args = parser.parse_args() + + # Load input file + try: + with open(args.input_file, 'r') as f: + input_data = json.load(f) + except FileNotFoundError: + print(f"Error: Input file '{args.input_file}' not found.", file=sys.stderr) + return 1 + except json.JSONDecodeError as e: + print(f"Error: Invalid JSON in '{args.input_file}': {e}", file=sys.stderr) + return 1 + + # Initialize linter and run analysis + linter = APILinter() + + try: + if args.raw_endpoints: + report = linter.lint_raw_endpoints(input_data) + else: + report = linter.lint_openapi_spec(input_data) + except Exception as e: + print(f"Error during linting: {e}", file=sys.stderr) + return 1 + + # Generate report + if args.format == 'json': + output = linter.generate_json_report() + else: + output = linter.generate_text_report() + + # Write output + if args.output: + try: + with open(args.output, 'w') as f: + f.write(output) + print(f"Report written to {args.output}") + except IOError as e: + print(f"Error writing to '{args.output}': {e}", file=sys.stderr) + return 1 + else: + print(output) + + # Return appropriate exit code + error_count = len([i for i in report.issues if i.severity == 'error']) + return 1 if error_count > 0 else 0 + + +if __name__ == '__main__': + sys.exit(main()) \ No newline at end of file diff --git a/skills/api-design-reviewer/scripts/api_scorecard.py b/skills/api-design-reviewer/scripts/api_scorecard.py new file mode 100644 index 00000000..dc673363 --- /dev/null +++ b/skills/api-design-reviewer/scripts/api_scorecard.py @@ -0,0 +1,1661 @@ +#!/usr/bin/env python3 +""" +API Scorecard - Comprehensive API design quality assessment tool. + +This script evaluates API designs across multiple dimensions and generates +a detailed scorecard with letter grades and improvement recommendations. + +Scoring Dimensions: +- Consistency (30%): Naming conventions, response patterns, structural consistency +- Documentation (20%): Completeness and clarity of API documentation +- Security (20%): Authentication, authorization, and security best practices +- Usability (15%): Ease of use, discoverability, and developer experience +- Performance (15%): Caching, pagination, and efficiency patterns + +Generates letter grades (A-F) with detailed breakdowns and actionable recommendations. +""" + +import argparse +import json +import re +import sys +from typing import Any, Dict, List, Optional, Set, Tuple +from dataclasses import dataclass, field +from enum import Enum +import math + + +class ScoreCategory(Enum): + """Scoring categories.""" + CONSISTENCY = "consistency" + DOCUMENTATION = "documentation" + SECURITY = "security" + USABILITY = "usability" + PERFORMANCE = "performance" + + +@dataclass +class CategoryScore: + """Score for a specific category.""" + category: ScoreCategory + score: float # 0-100 + max_score: float # Usually 100 + weight: float # Percentage weight in overall score + issues: List[str] = field(default_factory=list) + recommendations: List[str] = field(default_factory=list) + + @property + def letter_grade(self) -> str: + """Convert score to letter grade.""" + if self.score >= 90: + return "A" + elif self.score >= 80: + return "B" + elif self.score >= 70: + return "C" + elif self.score >= 60: + return "D" + else: + return "F" + + @property + def weighted_score(self) -> float: + """Calculate weighted contribution to overall score.""" + return (self.score / 100.0) * self.weight + + +@dataclass +class APIScorecard: + """Complete API scorecard with all category scores.""" + category_scores: Dict[ScoreCategory, CategoryScore] = field(default_factory=dict) + overall_score: float = 0.0 + overall_grade: str = "F" + total_endpoints: int = 0 + api_info: Dict[str, Any] = field(default_factory=dict) + + def calculate_overall_score(self) -> None: + """Calculate overall weighted score and grade.""" + self.overall_score = sum(score.weighted_score for score in self.category_scores.values()) + + if self.overall_score >= 90: + self.overall_grade = "A" + elif self.overall_score >= 80: + self.overall_grade = "B" + elif self.overall_score >= 70: + self.overall_grade = "C" + elif self.overall_score >= 60: + self.overall_grade = "D" + else: + self.overall_grade = "F" + + def get_top_recommendations(self, limit: int = 5) -> List[str]: + """Get top recommendations across all categories.""" + all_recommendations = [] + for category_score in self.category_scores.values(): + for rec in category_score.recommendations: + all_recommendations.append(f"{category_score.category.value.title()}: {rec}") + + # Sort by category weight (highest impact first) + weighted_recs = [] + for category_score in sorted(self.category_scores.values(), + key=lambda x: x.weight, reverse=True): + for rec in category_score.recommendations[:2]: # Top 2 per category + weighted_recs.append(f"{category_score.category.value.title()}: {rec}") + + return weighted_recs[:limit] + + +class APIScoringEngine: + """Main API scoring engine.""" + + def __init__(self): + self.scorecard = APIScorecard() + self.spec: Optional[Dict] = None + + # Regex patterns for validation + self.kebab_case_pattern = re.compile(r'^[a-z]+(?:-[a-z0-9]+)*$') + self.camel_case_pattern = re.compile(r'^[a-z][a-zA-Z0-9]*$') + self.pascal_case_pattern = re.compile(r'^[A-Z][a-zA-Z0-9]*$') + + # HTTP methods + self.http_methods = {'GET', 'POST', 'PUT', 'PATCH', 'DELETE', 'HEAD', 'OPTIONS'} + + # Category weights (must sum to 100) + self.category_weights = { + ScoreCategory.CONSISTENCY: 30.0, + ScoreCategory.DOCUMENTATION: 20.0, + ScoreCategory.SECURITY: 20.0, + ScoreCategory.USABILITY: 15.0, + ScoreCategory.PERFORMANCE: 15.0 + } + + def score_api(self, spec: Dict[str, Any]) -> APIScorecard: + """Generate comprehensive API scorecard.""" + self.spec = spec + self.scorecard = APIScorecard() + + # Extract basic API info + self._extract_api_info() + + # Score each category + self._score_consistency() + self._score_documentation() + self._score_security() + self._score_usability() + self._score_performance() + + # Calculate overall score + self.scorecard.calculate_overall_score() + + return self.scorecard + + def _extract_api_info(self) -> None: + """Extract basic API information.""" + info = self.spec.get('info', {}) + paths = self.spec.get('paths', {}) + + self.scorecard.api_info = { + 'title': info.get('title', 'Unknown API'), + 'version': info.get('version', ''), + 'description': info.get('description', ''), + 'total_paths': len(paths), + 'openapi_version': self.spec.get('openapi', self.spec.get('swagger', '')) + } + + # Count total endpoints + endpoint_count = 0 + for path_obj in paths.values(): + if isinstance(path_obj, dict): + endpoint_count += len([m for m in path_obj.keys() + if m.upper() in self.http_methods]) + + self.scorecard.total_endpoints = endpoint_count + + def _score_consistency(self) -> None: + """Score API consistency (30% weight).""" + category = ScoreCategory.CONSISTENCY + score = CategoryScore( + category=category, + score=0.0, + max_score=100.0, + weight=self.category_weights[category] + ) + + consistency_checks = [ + self._check_naming_consistency(), + self._check_response_consistency(), + self._check_error_format_consistency(), + self._check_parameter_consistency(), + self._check_url_structure_consistency(), + self._check_http_method_consistency(), + self._check_status_code_consistency() + ] + + # Average the consistency scores + valid_scores = [s for s in consistency_checks if s is not None] + if valid_scores: + score.score = sum(valid_scores) / len(valid_scores) + + # Add specific recommendations based on low scores + if score.score < 70: + score.recommendations.extend([ + "Review naming conventions across all endpoints and schemas", + "Standardize response formats and error structures", + "Ensure consistent HTTP method usage patterns" + ]) + elif score.score < 85: + score.recommendations.extend([ + "Minor consistency improvements needed in naming or response formats", + "Consider creating API design guidelines document" + ]) + + self.scorecard.category_scores[category] = score + + def _check_naming_consistency(self) -> float: + """Check naming convention consistency.""" + paths = self.spec.get('paths', {}) + schemas = self.spec.get('components', {}).get('schemas', {}) + + total_checks = 0 + passed_checks = 0 + + # Check path naming (should be kebab-case) + for path in paths.keys(): + segments = [seg for seg in path.split('/') if seg and not seg.startswith('{')] + for segment in segments: + total_checks += 1 + if self.kebab_case_pattern.match(segment) or re.match(r'^v\d+$', segment): + passed_checks += 1 + + # Check schema naming (should be PascalCase) + for schema_name in schemas.keys(): + total_checks += 1 + if self.pascal_case_pattern.match(schema_name): + passed_checks += 1 + + # Check property naming within schemas + for schema in schemas.values(): + if isinstance(schema, dict) and 'properties' in schema: + for prop_name in schema['properties'].keys(): + total_checks += 1 + if self.camel_case_pattern.match(prop_name): + passed_checks += 1 + + return (passed_checks / total_checks * 100) if total_checks > 0 else 100 + + def _check_response_consistency(self) -> float: + """Check response format consistency.""" + paths = self.spec.get('paths', {}) + + response_patterns = [] + total_responses = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods or not isinstance(operation, dict): + continue + + responses = operation.get('responses', {}) + for status_code, response in responses.items(): + if not isinstance(response, dict): + continue + + total_responses += 1 + content = response.get('content', {}) + + # Analyze response structure + for media_type, media_obj in content.items(): + schema = media_obj.get('schema', {}) + pattern = self._extract_schema_pattern(schema) + response_patterns.append(pattern) + + # Calculate consistency by comparing patterns + if not response_patterns: + return 100 + + pattern_counts = {} + for pattern in response_patterns: + pattern_key = json.dumps(pattern, sort_keys=True) + pattern_counts[pattern_key] = pattern_counts.get(pattern_key, 0) + 1 + + # Most common pattern should dominate for good consistency + max_count = max(pattern_counts.values()) if pattern_counts else 0 + consistency_ratio = max_count / len(response_patterns) if response_patterns else 1 + + return consistency_ratio * 100 + + def _extract_schema_pattern(self, schema: Dict[str, Any]) -> Dict[str, Any]: + """Extract a pattern from a schema for consistency checking.""" + if not isinstance(schema, dict): + return {} + + pattern = { + 'type': schema.get('type'), + 'has_properties': 'properties' in schema, + 'has_items': 'items' in schema, + 'required_count': len(schema.get('required', [])), + 'property_count': len(schema.get('properties', {})) + } + + return pattern + + def _check_error_format_consistency(self) -> float: + """Check error response format consistency.""" + paths = self.spec.get('paths', {}) + error_responses = [] + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + responses = operation.get('responses', {}) + for status_code, response in responses.items(): + try: + code_int = int(status_code) + if code_int >= 400: # Error responses + content = response.get('content', {}) + for media_type, media_obj in content.items(): + schema = media_obj.get('schema', {}) + error_responses.append(self._extract_schema_pattern(schema)) + except ValueError: + continue + + if not error_responses: + return 80 # No error responses defined - somewhat concerning + + # Check consistency of error response formats + pattern_counts = {} + for pattern in error_responses: + pattern_key = json.dumps(pattern, sort_keys=True) + pattern_counts[pattern_key] = pattern_counts.get(pattern_key, 0) + 1 + + max_count = max(pattern_counts.values()) if pattern_counts else 0 + consistency_ratio = max_count / len(error_responses) if error_responses else 1 + + return consistency_ratio * 100 + + def _check_parameter_consistency(self) -> float: + """Check parameter naming and usage consistency.""" + paths = self.spec.get('paths', {}) + + query_params = [] + path_params = [] + header_params = [] + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + parameters = operation.get('parameters', []) + for param in parameters: + if not isinstance(param, dict): + continue + + param_name = param.get('name', '') + param_in = param.get('in', '') + + if param_in == 'query': + query_params.append(param_name) + elif param_in == 'path': + path_params.append(param_name) + elif param_in == 'header': + header_params.append(param_name) + + # Check naming consistency for each parameter type + scores = [] + + # Query parameters should be camelCase or kebab-case + if query_params: + valid_query = sum(1 for p in query_params + if self.camel_case_pattern.match(p) or self.kebab_case_pattern.match(p)) + scores.append((valid_query / len(query_params)) * 100) + + # Path parameters should be camelCase or kebab-case + if path_params: + valid_path = sum(1 for p in path_params + if self.camel_case_pattern.match(p) or self.kebab_case_pattern.match(p)) + scores.append((valid_path / len(path_params)) * 100) + + return sum(scores) / len(scores) if scores else 100 + + def _check_url_structure_consistency(self) -> float: + """Check URL structure and pattern consistency.""" + paths = self.spec.get('paths', {}) + + total_paths = len(paths) + if total_paths == 0: + return 0 + + structure_score = 0 + + # Check for consistent versioning + versioned_paths = 0 + for path in paths.keys(): + if re.search(r'/v\d+/', path): + versioned_paths += 1 + + # Either all or none should be versioned for consistency + if versioned_paths == 0 or versioned_paths == total_paths: + structure_score += 25 + elif versioned_paths > total_paths * 0.8: + structure_score += 20 + + # Check for reasonable path depth + reasonable_depth = 0 + for path in paths.keys(): + segments = [seg for seg in path.split('/') if seg] + if 2 <= len(segments) <= 5: # Reasonable depth + reasonable_depth += 1 + + structure_score += (reasonable_depth / total_paths) * 25 + + # Check for RESTful resource patterns + restful_patterns = 0 + for path in paths.keys(): + # Look for patterns like /resources/{id} or /resources + if re.match(r'^/[a-z-]+(/\{[^}]+\})?(/[a-z-]+)*$', path): + restful_patterns += 1 + + structure_score += (restful_patterns / total_paths) * 30 + + # Check for consistent trailing slash usage + with_slash = sum(1 for path in paths.keys() if path.endswith('/')) + without_slash = total_paths - with_slash + + # Either all or none should have trailing slashes + if with_slash == 0 or without_slash == 0: + structure_score += 20 + elif min(with_slash, without_slash) < total_paths * 0.1: + structure_score += 15 + + return min(structure_score, 100) + + def _check_http_method_consistency(self) -> float: + """Check HTTP method usage consistency.""" + paths = self.spec.get('paths', {}) + + method_usage = {} + total_operations = 0 + + for path, path_obj in paths.items(): + if not isinstance(path_obj, dict): + continue + + for method in path_obj.keys(): + if method.upper() in self.http_methods: + method_upper = method.upper() + total_operations += 1 + + # Analyze method usage patterns + if method_upper not in method_usage: + method_usage[method_upper] = {'count': 0, 'appropriate': 0} + + method_usage[method_upper]['count'] += 1 + + # Check if method usage seems appropriate + if self._is_method_usage_appropriate(path, method_upper, path_obj[method]): + method_usage[method_upper]['appropriate'] += 1 + + if total_operations == 0: + return 0 + + # Calculate appropriateness score + total_appropriate = sum(data['appropriate'] for data in method_usage.values()) + return (total_appropriate / total_operations) * 100 + + def _is_method_usage_appropriate(self, path: str, method: str, operation: Dict) -> bool: + """Check if HTTP method usage is appropriate for the endpoint.""" + # Simple heuristics for method appropriateness + has_request_body = 'requestBody' in operation + path_has_id = '{' in path and '}' in path + + if method == 'GET': + return not has_request_body # GET should not have body + elif method == 'POST': + return not path_has_id # POST typically for collections + elif method == 'PUT': + return path_has_id and has_request_body # PUT for specific resources + elif method == 'PATCH': + return path_has_id # PATCH for specific resources + elif method == 'DELETE': + return path_has_id # DELETE for specific resources + + return True # Default to appropriate for other methods + + def _check_status_code_consistency(self) -> float: + """Check HTTP status code usage consistency.""" + paths = self.spec.get('paths', {}) + + method_status_patterns = {} + total_operations = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + responses = operation.get('responses', {}) + status_codes = set(responses.keys()) + + if method.upper() not in method_status_patterns: + method_status_patterns[method.upper()] = [] + + method_status_patterns[method.upper()].append(status_codes) + + if total_operations == 0: + return 0 + + # Check consistency within each method type + consistency_scores = [] + + for method, status_patterns in method_status_patterns.items(): + if not status_patterns: + continue + + # Find common status codes for this method + all_codes = set() + for pattern in status_patterns: + all_codes.update(pattern) + + # Calculate how many operations use the most common codes + code_usage = {} + for code in all_codes: + code_usage[code] = sum(1 for pattern in status_patterns if code in pattern) + + # Score based on consistency of common status codes + if status_patterns: + avg_consistency = sum( + len([code for code in pattern if code_usage.get(code, 0) > len(status_patterns) * 0.5]) + for pattern in status_patterns + ) / len(status_patterns) + + method_consistency = min(avg_consistency / 3.0 * 100, 100) # Expect ~3 common codes + consistency_scores.append(method_consistency) + + return sum(consistency_scores) / len(consistency_scores) if consistency_scores else 100 + + def _score_documentation(self) -> None: + """Score API documentation quality (20% weight).""" + category = ScoreCategory.DOCUMENTATION + score = CategoryScore( + category=category, + score=0.0, + max_score=100.0, + weight=self.category_weights[category] + ) + + documentation_checks = [ + self._check_api_level_documentation(), + self._check_endpoint_documentation(), + self._check_schema_documentation(), + self._check_parameter_documentation(), + self._check_response_documentation(), + self._check_example_coverage() + ] + + valid_scores = [s for s in documentation_checks if s is not None] + if valid_scores: + score.score = sum(valid_scores) / len(valid_scores) + + # Add recommendations based on score + if score.score < 60: + score.recommendations.extend([ + "Add comprehensive descriptions to all API components", + "Include examples for complex operations and schemas", + "Document all parameters and response fields" + ]) + elif score.score < 80: + score.recommendations.extend([ + "Improve documentation completeness for some endpoints", + "Add more examples to enhance developer experience" + ]) + + self.scorecard.category_scores[category] = score + + def _check_api_level_documentation(self) -> float: + """Check API-level documentation completeness.""" + info = self.spec.get('info', {}) + score = 0 + + # Required fields + if info.get('title'): + score += 20 + if info.get('version'): + score += 20 + if info.get('description') and len(info['description']) > 20: + score += 30 + + # Optional but recommended fields + if info.get('contact'): + score += 15 + if info.get('license'): + score += 15 + + return score + + def _check_endpoint_documentation(self) -> float: + """Check endpoint-level documentation completeness.""" + paths = self.spec.get('paths', {}) + + total_operations = 0 + documented_operations = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + doc_score = 0 + + if operation.get('summary'): + doc_score += 1 + if operation.get('description') and len(operation['description']) > 20: + doc_score += 1 + if operation.get('operationId'): + doc_score += 1 + + # Consider it documented if it has at least 2/3 elements + if doc_score >= 2: + documented_operations += 1 + + return (documented_operations / total_operations * 100) if total_operations > 0 else 100 + + def _check_schema_documentation(self) -> float: + """Check schema documentation completeness.""" + schemas = self.spec.get('components', {}).get('schemas', {}) + + if not schemas: + return 80 # No schemas to document + + total_schemas = len(schemas) + documented_schemas = 0 + + for schema_name, schema in schemas.items(): + if not isinstance(schema, dict): + continue + + doc_elements = 0 + + # Schema-level description + if schema.get('description'): + doc_elements += 1 + + # Property descriptions + properties = schema.get('properties', {}) + if properties: + described_props = sum(1 for prop in properties.values() + if isinstance(prop, dict) and prop.get('description')) + if described_props > len(properties) * 0.5: # At least 50% documented + doc_elements += 1 + + # Examples + if schema.get('example') or any( + isinstance(prop, dict) and prop.get('example') + for prop in properties.values() + ): + doc_elements += 1 + + if doc_elements >= 2: + documented_schemas += 1 + + return (documented_schemas / total_schemas * 100) if total_schemas > 0 else 100 + + def _check_parameter_documentation(self) -> float: + """Check parameter documentation completeness.""" + paths = self.spec.get('paths', {}) + + total_params = 0 + documented_params = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + parameters = operation.get('parameters', []) + for param in parameters: + if not isinstance(param, dict): + continue + + total_params += 1 + + doc_score = 0 + if param.get('description'): + doc_score += 1 + if param.get('example') or (param.get('schema', {}).get('example')): + doc_score += 1 + + if doc_score >= 1: # At least description + documented_params += 1 + + return (documented_params / total_params * 100) if total_params > 0 else 100 + + def _check_response_documentation(self) -> float: + """Check response documentation completeness.""" + paths = self.spec.get('paths', {}) + + total_responses = 0 + documented_responses = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + responses = operation.get('responses', {}) + for status_code, response in responses.items(): + if not isinstance(response, dict): + continue + + total_responses += 1 + + if response.get('description'): + documented_responses += 1 + + return (documented_responses / total_responses * 100) if total_responses > 0 else 100 + + def _check_example_coverage(self) -> float: + """Check example coverage across the API.""" + paths = self.spec.get('paths', {}) + schemas = self.spec.get('components', {}).get('schemas', {}) + + # Check examples in operations + total_operations = 0 + operations_with_examples = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + + has_example = False + + # Check request body examples + request_body = operation.get('requestBody', {}) + if self._has_examples(request_body.get('content', {})): + has_example = True + + # Check response examples + responses = operation.get('responses', {}) + for response in responses.values(): + if isinstance(response, dict) and self._has_examples(response.get('content', {})): + has_example = True + break + + if has_example: + operations_with_examples += 1 + + # Check examples in schemas + total_schemas = len(schemas) + schemas_with_examples = 0 + + for schema in schemas.values(): + if isinstance(schema, dict) and self._schema_has_examples(schema): + schemas_with_examples += 1 + + # Combine scores + operation_score = (operations_with_examples / total_operations * 100) if total_operations > 0 else 100 + schema_score = (schemas_with_examples / total_schemas * 100) if total_schemas > 0 else 100 + + return (operation_score + schema_score) / 2 + + def _has_examples(self, content: Dict[str, Any]) -> bool: + """Check if content has examples.""" + for media_type, media_obj in content.items(): + if isinstance(media_obj, dict): + if media_obj.get('example') or media_obj.get('examples'): + return True + return False + + def _schema_has_examples(self, schema: Dict[str, Any]) -> bool: + """Check if schema has examples.""" + if schema.get('example'): + return True + + properties = schema.get('properties', {}) + for prop in properties.values(): + if isinstance(prop, dict) and prop.get('example'): + return True + + return False + + def _score_security(self) -> None: + """Score API security implementation (20% weight).""" + category = ScoreCategory.SECURITY + score = CategoryScore( + category=category, + score=0.0, + max_score=100.0, + weight=self.category_weights[category] + ) + + security_checks = [ + self._check_security_schemes(), + self._check_security_requirements(), + self._check_https_usage(), + self._check_authentication_patterns(), + self._check_sensitive_data_handling() + ] + + valid_scores = [s for s in security_checks if s is not None] + if valid_scores: + score.score = sum(valid_scores) / len(valid_scores) + + # Add recommendations + if score.score < 50: + score.recommendations.extend([ + "Implement comprehensive security schemes (OAuth2, API keys, etc.)", + "Ensure all endpoints have appropriate security requirements", + "Add input validation and rate limiting patterns" + ]) + elif score.score < 80: + score.recommendations.extend([ + "Review security coverage for all endpoints", + "Consider additional security measures for sensitive operations" + ]) + + self.scorecard.category_scores[category] = score + + def _check_security_schemes(self) -> float: + """Check security scheme definitions.""" + security_schemes = self.spec.get('components', {}).get('securitySchemes', {}) + + if not security_schemes: + return 20 # Very low score for no security + + score = 40 # Base score for having security schemes + + scheme_types = set() + for scheme in security_schemes.values(): + if isinstance(scheme, dict): + scheme_type = scheme.get('type') + scheme_types.add(scheme_type) + + # Bonus for modern security schemes + if 'oauth2' in scheme_types: + score += 30 + if 'apiKey' in scheme_types: + score += 15 + if 'http' in scheme_types: + score += 15 + + return min(score, 100) + + def _check_security_requirements(self) -> float: + """Check security requirement coverage.""" + paths = self.spec.get('paths', {}) + global_security = self.spec.get('security', []) + + total_operations = 0 + secured_operations = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + + # Check if operation has security requirements + operation_security = operation.get('security') + + if operation_security is not None: + secured_operations += 1 + elif global_security: + secured_operations += 1 + + return (secured_operations / total_operations * 100) if total_operations > 0 else 0 + + def _check_https_usage(self) -> float: + """Check HTTPS enforcement.""" + servers = self.spec.get('servers', []) + + if not servers: + return 60 # No servers defined - assume HTTPS + + https_servers = 0 + for server in servers: + if isinstance(server, dict): + url = server.get('url', '') + if url.startswith('https://') or not url.startswith('http://'): + https_servers += 1 + + return (https_servers / len(servers) * 100) if servers else 100 + + def _check_authentication_patterns(self) -> float: + """Check authentication pattern quality.""" + security_schemes = self.spec.get('components', {}).get('securitySchemes', {}) + + if not security_schemes: + return 0 + + pattern_scores = [] + + for scheme in security_schemes.values(): + if not isinstance(scheme, dict): + continue + + scheme_type = scheme.get('type', '').lower() + + if scheme_type == 'oauth2': + # OAuth2 is highly recommended + flows = scheme.get('flows', {}) + if flows: + pattern_scores.append(95) + else: + pattern_scores.append(80) + elif scheme_type == 'http': + scheme_scheme = scheme.get('scheme', '').lower() + if scheme_scheme == 'bearer': + pattern_scores.append(85) + elif scheme_scheme == 'basic': + pattern_scores.append(60) # Less secure + else: + pattern_scores.append(70) + elif scheme_type == 'apikey': + location = scheme.get('in', '').lower() + if location == 'header': + pattern_scores.append(75) + else: + pattern_scores.append(60) # Query/cookie less secure + else: + pattern_scores.append(50) # Unknown scheme + + return sum(pattern_scores) / len(pattern_scores) if pattern_scores else 0 + + def _check_sensitive_data_handling(self) -> float: + """Check sensitive data handling patterns.""" + # This is a simplified check - in reality would need more sophisticated analysis + schemas = self.spec.get('components', {}).get('schemas', {}) + + score = 80 # Default good score + + # Look for potential sensitive fields without proper handling + sensitive_field_names = {'password', 'secret', 'token', 'key', 'ssn', 'credit_card'} + + for schema in schemas.values(): + if not isinstance(schema, dict): + continue + + properties = schema.get('properties', {}) + for prop_name, prop_def in properties.items(): + if not isinstance(prop_def, dict): + continue + + # Check for sensitive field names + if any(sensitive in prop_name.lower() for sensitive in sensitive_field_names): + # Check if it's marked as sensitive (writeOnly, format: password, etc.) + if not (prop_def.get('writeOnly') or + prop_def.get('format') == 'password' or + 'password' in prop_def.get('description', '').lower()): + score -= 10 # Penalty for exposed sensitive field + + return max(score, 0) + + def _score_usability(self) -> None: + """Score API usability and developer experience (15% weight).""" + category = ScoreCategory.USABILITY + score = CategoryScore( + category=category, + score=0.0, + max_score=100.0, + weight=self.category_weights[category] + ) + + usability_checks = [ + self._check_discoverability(), + self._check_error_handling(), + self._check_filtering_and_searching(), + self._check_resource_relationships(), + self._check_developer_experience() + ] + + valid_scores = [s for s in usability_checks if s is not None] + if valid_scores: + score.score = sum(valid_scores) / len(valid_scores) + + # Add recommendations + if score.score < 60: + score.recommendations.extend([ + "Improve error messages with actionable guidance", + "Add filtering and search capabilities to list endpoints", + "Enhance resource discoverability with better linking" + ]) + elif score.score < 80: + score.recommendations.extend([ + "Consider adding HATEOAS links for better discoverability", + "Enhance developer experience with better examples" + ]) + + self.scorecard.category_scores[category] = score + + def _check_discoverability(self) -> float: + """Check API discoverability features.""" + paths = self.spec.get('paths', {}) + + # Look for root/discovery endpoints + has_root = '/' in paths or any(path == '/api' or path.startswith('/api/') for path in paths) + + # Look for HATEOAS patterns in responses + hateoas_score = 0 + total_responses = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + responses = operation.get('responses', {}) + for response in responses.values(): + if not isinstance(response, dict): + continue + + total_responses += 1 + + # Look for link-like properties in response schemas + content = response.get('content', {}) + for media_obj in content.values(): + schema = media_obj.get('schema', {}) + if self._has_link_properties(schema): + hateoas_score += 1 + break + + discovery_score = 50 if has_root else 30 + if total_responses > 0: + hateoas_ratio = hateoas_score / total_responses + discovery_score += hateoas_ratio * 50 + + return min(discovery_score, 100) + + def _has_link_properties(self, schema: Dict[str, Any]) -> bool: + """Check if schema has link-like properties.""" + if not isinstance(schema, dict): + return False + + properties = schema.get('properties', {}) + link_indicators = {'links', '_links', 'href', 'url', 'self', 'next', 'prev'} + + return any(prop_name.lower() in link_indicators for prop_name in properties.keys()) + + def _check_error_handling(self) -> float: + """Check error handling quality.""" + paths = self.spec.get('paths', {}) + + total_operations = 0 + operations_with_errors = 0 + detailed_error_responses = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + responses = operation.get('responses', {}) + + # Check for error responses + has_error_responses = any( + status_code.startswith('4') or status_code.startswith('5') + for status_code in responses.keys() + ) + + if has_error_responses: + operations_with_errors += 1 + + # Check for detailed error schemas + for status_code, response in responses.items(): + if (status_code.startswith('4') or status_code.startswith('5')) and isinstance(response, dict): + content = response.get('content', {}) + for media_obj in content.values(): + schema = media_obj.get('schema', {}) + if self._has_detailed_error_schema(schema): + detailed_error_responses += 1 + break + break + + if total_operations == 0: + return 0 + + error_coverage = (operations_with_errors / total_operations) * 60 + error_detail = (detailed_error_responses / operations_with_errors * 40) if operations_with_errors > 0 else 0 + + return error_coverage + error_detail + + def _has_detailed_error_schema(self, schema: Dict[str, Any]) -> bool: + """Check if error schema has detailed information.""" + if not isinstance(schema, dict): + return False + + properties = schema.get('properties', {}) + error_fields = {'error', 'message', 'details', 'code', 'timestamp'} + + matching_fields = sum(1 for field in error_fields if field in properties) + return matching_fields >= 2 # At least 2 standard error fields + + def _check_filtering_and_searching(self) -> float: + """Check filtering and search capabilities.""" + paths = self.spec.get('paths', {}) + + collection_endpoints = 0 + endpoints_with_filtering = 0 + + for path, path_obj in paths.items(): + if not isinstance(path_obj, dict): + continue + + # Identify collection endpoints (no path parameters) + if '{' not in path: + get_operation = path_obj.get('get') + if get_operation: + collection_endpoints += 1 + + # Check for filtering/search parameters + parameters = get_operation.get('parameters', []) + filter_params = {'filter', 'search', 'q', 'query', 'limit', 'page', 'offset'} + + has_filtering = any( + isinstance(param, dict) and param.get('name', '').lower() in filter_params + for param in parameters + ) + + if has_filtering: + endpoints_with_filtering += 1 + + return (endpoints_with_filtering / collection_endpoints * 100) if collection_endpoints > 0 else 100 + + def _check_resource_relationships(self) -> float: + """Check resource relationship handling.""" + paths = self.spec.get('paths', {}) + schemas = self.spec.get('components', {}).get('schemas', {}) + + # Look for nested resource patterns + nested_resources = 0 + total_resource_paths = 0 + + for path in paths.keys(): + # Skip root paths + if path.count('/') >= 3: # e.g., /api/users/123/orders + total_resource_paths += 1 + if '{' in path: + nested_resources += 1 + + # Look for relationship fields in schemas + schemas_with_relations = 0 + for schema in schemas.values(): + if not isinstance(schema, dict): + continue + + properties = schema.get('properties', {}) + relation_indicators = {'id', '_id', 'ref', 'link', 'relationship'} + + has_relations = any( + any(indicator in prop_name.lower() for indicator in relation_indicators) + for prop_name in properties.keys() + ) + + if has_relations: + schemas_with_relations += 1 + + nested_score = (nested_resources / total_resource_paths * 50) if total_resource_paths > 0 else 25 + schema_score = (schemas_with_relations / len(schemas) * 50) if schemas else 25 + + return nested_score + schema_score + + def _check_developer_experience(self) -> float: + """Check overall developer experience factors.""" + # This is a composite score based on various DX factors + factors = [] + + # Factor 1: Consistent response structure + factors.append(self._check_response_consistency()) + + # Factor 2: Clear operation IDs + paths = self.spec.get('paths', {}) + total_operations = 0 + operations_with_ids = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method, operation in path_obj.items(): + if method.upper() not in self.http_methods: + continue + + total_operations += 1 + if isinstance(operation, dict) and operation.get('operationId'): + operations_with_ids += 1 + + operation_id_score = (operations_with_ids / total_operations * 100) if total_operations > 0 else 100 + factors.append(operation_id_score) + + # Factor 3: Reasonable path complexity + avg_path_complexity = 0 + if paths: + complexities = [] + for path in paths.keys(): + segments = [seg for seg in path.split('/') if seg] + complexities.append(len(segments)) + + avg_complexity = sum(complexities) / len(complexities) + # Optimal complexity is 3-4 segments + if 3 <= avg_complexity <= 4: + avg_path_complexity = 100 + elif 2 <= avg_complexity <= 5: + avg_path_complexity = 80 + else: + avg_path_complexity = 60 + + factors.append(avg_path_complexity) + + return sum(factors) / len(factors) if factors else 0 + + def _score_performance(self) -> None: + """Score API performance patterns (15% weight).""" + category = ScoreCategory.PERFORMANCE + score = CategoryScore( + category=category, + score=0.0, + max_score=100.0, + weight=self.category_weights[category] + ) + + performance_checks = [ + self._check_caching_headers(), + self._check_pagination_patterns(), + self._check_compression_support(), + self._check_efficiency_patterns(), + self._check_batch_operations() + ] + + valid_scores = [s for s in performance_checks if s is not None] + if valid_scores: + score.score = sum(valid_scores) / len(valid_scores) + + # Add recommendations + if score.score < 60: + score.recommendations.extend([ + "Implement pagination for list endpoints", + "Add caching headers for cacheable responses", + "Consider batch operations for bulk updates" + ]) + elif score.score < 80: + score.recommendations.extend([ + "Review caching strategies for better performance", + "Consider field selection parameters for large responses" + ]) + + self.scorecard.category_scores[category] = score + + def _check_caching_headers(self) -> float: + """Check caching header implementation.""" + paths = self.spec.get('paths', {}) + + get_operations = 0 + cacheable_operations = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + get_operation = path_obj.get('get') + if get_operation and isinstance(get_operation, dict): + get_operations += 1 + + # Check for caching-related headers in responses + responses = get_operation.get('responses', {}) + for response in responses.values(): + if not isinstance(response, dict): + continue + + headers = response.get('headers', {}) + cache_headers = {'cache-control', 'etag', 'last-modified', 'expires'} + + if any(header.lower() in cache_headers for header in headers.keys()): + cacheable_operations += 1 + break + + return (cacheable_operations / get_operations * 100) if get_operations > 0 else 50 + + def _check_pagination_patterns(self) -> float: + """Check pagination implementation.""" + paths = self.spec.get('paths', {}) + + collection_endpoints = 0 + paginated_endpoints = 0 + + for path, path_obj in paths.items(): + if not isinstance(path_obj, dict): + continue + + # Identify collection endpoints + if '{' not in path: # No path parameters = collection + get_operation = path_obj.get('get') + if get_operation and isinstance(get_operation, dict): + collection_endpoints += 1 + + # Check for pagination parameters + parameters = get_operation.get('parameters', []) + pagination_params = {'limit', 'offset', 'page', 'pagesize', 'per_page', 'cursor'} + + has_pagination = any( + isinstance(param, dict) and param.get('name', '').lower() in pagination_params + for param in parameters + ) + + if has_pagination: + paginated_endpoints += 1 + + return (paginated_endpoints / collection_endpoints * 100) if collection_endpoints > 0 else 100 + + def _check_compression_support(self) -> float: + """Check compression support indicators.""" + # This is speculative - OpenAPI doesn't directly specify compression + # Look for indicators that compression is considered + + servers = self.spec.get('servers', []) + + # Check if any server descriptions mention compression + compression_mentions = 0 + for server in servers: + if isinstance(server, dict): + description = server.get('description', '').lower() + if any(term in description for term in ['gzip', 'compress', 'deflate']): + compression_mentions += 1 + + # Base score - assume compression is handled at server level + base_score = 70 + + if compression_mentions > 0: + return min(base_score + (compression_mentions * 10), 100) + + return base_score + + def _check_efficiency_patterns(self) -> float: + """Check efficiency patterns like field selection.""" + paths = self.spec.get('paths', {}) + + total_get_operations = 0 + operations_with_selection = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + get_operation = path_obj.get('get') + if get_operation and isinstance(get_operation, dict): + total_get_operations += 1 + + # Check for field selection parameters + parameters = get_operation.get('parameters', []) + selection_params = {'fields', 'select', 'include', 'exclude'} + + has_selection = any( + isinstance(param, dict) and param.get('name', '').lower() in selection_params + for param in parameters + ) + + if has_selection: + operations_with_selection += 1 + + return (operations_with_selection / total_get_operations * 100) if total_get_operations > 0 else 60 + + def _check_batch_operations(self) -> float: + """Check for batch operation support.""" + paths = self.spec.get('paths', {}) + + # Look for batch endpoints + batch_indicators = ['batch', 'bulk', 'multi'] + batch_endpoints = 0 + + for path in paths.keys(): + if any(indicator in path.lower() for indicator in batch_indicators): + batch_endpoints += 1 + + # Look for array-based request bodies (indicating batch operations) + array_operations = 0 + total_post_put_operations = 0 + + for path_obj in paths.values(): + if not isinstance(path_obj, dict): + continue + + for method in ['post', 'put', 'patch']: + operation = path_obj.get(method) + if operation and isinstance(operation, dict): + total_post_put_operations += 1 + + request_body = operation.get('requestBody', {}) + content = request_body.get('content', {}) + + for media_obj in content.values(): + schema = media_obj.get('schema', {}) + if schema.get('type') == 'array': + array_operations += 1 + break + + # Score based on presence of batch patterns + batch_score = min(batch_endpoints * 20, 60) # Up to 60 points for explicit batch endpoints + + if total_post_put_operations > 0: + array_score = (array_operations / total_post_put_operations) * 40 + batch_score += array_score + + return min(batch_score, 100) + + def generate_json_report(self) -> str: + """Generate JSON format scorecard.""" + report_data = { + "overall": { + "score": round(self.scorecard.overall_score, 2), + "grade": self.scorecard.overall_grade, + "totalEndpoints": self.scorecard.total_endpoints + }, + "api_info": self.scorecard.api_info, + "categories": {}, + "topRecommendations": self.scorecard.get_top_recommendations() + } + + for category, score in self.scorecard.category_scores.items(): + report_data["categories"][category.value] = { + "score": round(score.score, 2), + "grade": score.letter_grade, + "weight": score.weight, + "weightedScore": round(score.weighted_score, 2), + "issues": score.issues, + "recommendations": score.recommendations + } + + return json.dumps(report_data, indent=2) + + def generate_text_report(self) -> str: + """Generate human-readable scorecard report.""" + lines = [ + "═══════════════════════════════════════════════════════════════", + " API DESIGN SCORECARD", + "═══════════════════════════════════════════════════════════════", + f"API: {self.scorecard.api_info.get('title', 'Unknown')}", + f"Version: {self.scorecard.api_info.get('version', 'Unknown')}", + f"Total Endpoints: {self.scorecard.total_endpoints}", + "", + f"🏆 OVERALL GRADE: {self.scorecard.overall_grade} ({self.scorecard.overall_score:.1f}/100.0)", + "", + "═══════════════════════════════════════════════════════════════", + "DETAILED BREAKDOWN:", + "═══════════════════════════════════════════════════════════════" + ] + + # Sort categories by weight (most important first) + sorted_categories = sorted( + self.scorecard.category_scores.items(), + key=lambda x: x[1].weight, + reverse=True + ) + + for category, score in sorted_categories: + category_name = category.value.title().replace('_', ' ') + + lines.extend([ + "", + f"📊 {category_name.upper()} - Grade: {score.letter_grade} ({score.score:.1f}/100)", + f" Weight: {score.weight}% | Contribution: {score.weighted_score:.1f} points", + " " + "─" * 50 + ]) + + if score.recommendations: + lines.append(" 💡 Recommendations:") + for rec in score.recommendations[:3]: # Top 3 recommendations + lines.append(f" • {rec}") + else: + lines.append(" ✅ No specific recommendations - performing well!") + + # Overall assessment + lines.extend([ + "", + "═══════════════════════════════════════════════════════════════", + "OVERALL ASSESSMENT:", + "═══════════════════════════════════════════════════════════════" + ]) + + if self.scorecard.overall_grade == "A": + lines.extend([ + "🏆 EXCELLENT! Your API demonstrates outstanding design quality.", + " Continue following these best practices and consider sharing", + " your approach as a reference for other teams." + ]) + elif self.scorecard.overall_grade == "B": + lines.extend([ + "✅ GOOD! Your API follows most best practices with room for", + " minor improvements. Focus on the recommendations above", + " to achieve excellence." + ]) + elif self.scorecard.overall_grade == "C": + lines.extend([ + "⚠️ FAIR! Your API has a solid foundation but several areas", + " need improvement. Prioritize the high-weight categories", + " for maximum impact." + ]) + elif self.scorecard.overall_grade == "D": + lines.extend([ + "❌ NEEDS IMPROVEMENT! Your API has significant issues that", + " may impact developer experience and maintainability.", + " Focus on consistency and documentation first." + ]) + else: # Grade F + lines.extend([ + "🚨 CRITICAL ISSUES! Your API requires major redesign to meet", + " basic quality standards. Consider comprehensive review", + " of design principles and best practices." + ]) + + # Top recommendations + top_recs = self.scorecard.get_top_recommendations(3) + if top_recs: + lines.extend([ + "", + "🎯 TOP PRIORITY RECOMMENDATIONS:", + "" + ]) + for i, rec in enumerate(top_recs, 1): + lines.append(f" {i}. {rec}") + + lines.extend([ + "", + "═══════════════════════════════════════════════════════════════", + f"Generated by API Scorecard Tool | Score: {self.scorecard.overall_grade} ({self.scorecard.overall_score:.1f}%)", + "═══════════════════════════════════════════════════════════════" + ]) + + return "\n".join(lines) + + +def main(): + """Main CLI entry point.""" + parser = argparse.ArgumentParser( + description="Generate comprehensive API design quality scorecard", + formatter_class=argparse.RawDescriptionHelpFormatter, + epilog=""" +Examples: + python api_scorecard.py openapi.json + python api_scorecard.py --format json openapi.json > scorecard.json + python api_scorecard.py --output scorecard.txt openapi.json + """ + ) + + parser.add_argument( + 'spec_file', + help='OpenAPI/Swagger specification file (JSON format)' + ) + + parser.add_argument( + '--format', + choices=['text', 'json'], + default='text', + help='Output format (default: text)' + ) + + parser.add_argument( + '--output', + help='Output file (default: stdout)' + ) + + parser.add_argument( + '--min-grade', + choices=['A', 'B', 'C', 'D', 'F'], + help='Exit with code 1 if grade is below minimum' + ) + + args = parser.parse_args() + + # Load specification file + try: + with open(args.spec_file, 'r') as f: + spec = json.load(f) + except FileNotFoundError: + print(f"Error: Specification file '{args.spec_file}' not found.", file=sys.stderr) + return 1 + except json.JSONDecodeError as e: + print(f"Error: Invalid JSON in '{args.spec_file}': {e}", file=sys.stderr) + return 1 + + # Initialize scoring engine and generate scorecard + engine = APIScoringEngine() + + try: + scorecard = engine.score_api(spec) + except Exception as e: + print(f"Error during scoring: {e}", file=sys.stderr) + return 1 + + # Generate report + if args.format == 'json': + output = engine.generate_json_report() + else: + output = engine.generate_text_report() + + # Write output + if args.output: + try: + with open(args.output, 'w') as f: + f.write(output) + print(f"Scorecard written to {args.output}") + except IOError as e: + print(f"Error writing to '{args.output}': {e}", file=sys.stderr) + return 1 + else: + print(output) + + # Check minimum grade requirement + if args.min_grade: + grade_order = ['F', 'D', 'C', 'B', 'A'] + current_grade_index = grade_order.index(scorecard.overall_grade) + min_grade_index = grade_order.index(args.min_grade) + + if current_grade_index < min_grade_index: + print(f"Grade {scorecard.overall_grade} is below minimum required grade {args.min_grade}", file=sys.stderr) + return 1 + + return 0 + + +if __name__ == '__main__': + sys.exit(main()) \ No newline at end of file diff --git a/skills/api-design-reviewer/scripts/breaking_change_detector.py b/skills/api-design-reviewer/scripts/breaking_change_detector.py new file mode 100644 index 00000000..6f2736a9 --- /dev/null +++ b/skills/api-design-reviewer/scripts/breaking_change_detector.py @@ -0,0 +1,1102 @@ +#!/usr/bin/env python3 +""" +Breaking Change Detector - Compares API specification versions to identify breaking changes. + +This script analyzes two versions of an API specification and detects potentially +breaking changes including: +- Removed endpoints +- Modified response structures +- Removed or renamed fields +- Field type changes +- New required fields +- HTTP status code changes +- Parameter changes + +Generates detailed reports with migration guides for each breaking change. +""" + +import argparse +import json +import sys +from typing import Any, Dict, List, Set, Optional, Tuple, Union +from dataclasses import dataclass, field +from enum import Enum + + +class ChangeType(Enum): + """Types of API changes.""" + BREAKING = "breaking" + POTENTIALLY_BREAKING = "potentially_breaking" + NON_BREAKING = "non_breaking" + ENHANCEMENT = "enhancement" + + +class ChangeSeverity(Enum): + """Severity levels for changes.""" + CRITICAL = "critical" # Will definitely break clients + HIGH = "high" # Likely to break some clients + MEDIUM = "medium" # May break clients depending on usage + LOW = "low" # Minor impact, unlikely to break clients + INFO = "info" # Informational, no breaking impact + + +@dataclass +class Change: + """Represents a detected change between API versions.""" + change_type: ChangeType + severity: ChangeSeverity + category: str + path: str + message: str + old_value: Any = None + new_value: Any = None + migration_guide: str = "" + impact_description: str = "" + + def to_dict(self) -> Dict[str, Any]: + """Convert change to dictionary for JSON serialization.""" + return { + "changeType": self.change_type.value, + "severity": self.severity.value, + "category": self.category, + "path": self.path, + "message": self.message, + "oldValue": self.old_value, + "newValue": self.new_value, + "migrationGuide": self.migration_guide, + "impactDescription": self.impact_description + } + + +@dataclass +class ComparisonReport: + """Complete comparison report between two API versions.""" + changes: List[Change] = field(default_factory=list) + summary: Dict[str, int] = field(default_factory=dict) + + def add_change(self, change: Change) -> None: + """Add a change to the report.""" + self.changes.append(change) + + def calculate_summary(self) -> None: + """Calculate summary statistics.""" + self.summary = { + "total_changes": len(self.changes), + "breaking_changes": len([c for c in self.changes if c.change_type == ChangeType.BREAKING]), + "potentially_breaking_changes": len([c for c in self.changes if c.change_type == ChangeType.POTENTIALLY_BREAKING]), + "non_breaking_changes": len([c for c in self.changes if c.change_type == ChangeType.NON_BREAKING]), + "enhancements": len([c for c in self.changes if c.change_type == ChangeType.ENHANCEMENT]), + "critical_severity": len([c for c in self.changes if c.severity == ChangeSeverity.CRITICAL]), + "high_severity": len([c for c in self.changes if c.severity == ChangeSeverity.HIGH]), + "medium_severity": len([c for c in self.changes if c.severity == ChangeSeverity.MEDIUM]), + "low_severity": len([c for c in self.changes if c.severity == ChangeSeverity.LOW]), + "info_severity": len([c for c in self.changes if c.severity == ChangeSeverity.INFO]) + } + + def has_breaking_changes(self) -> bool: + """Check if report contains any breaking changes.""" + return any(c.change_type in [ChangeType.BREAKING, ChangeType.POTENTIALLY_BREAKING] + for c in self.changes) + + +class BreakingChangeDetector: + """Main breaking change detection engine.""" + + def __init__(self): + self.report = ComparisonReport() + self.old_spec: Optional[Dict] = None + self.new_spec: Optional[Dict] = None + + def compare_specs(self, old_spec: Dict[str, Any], new_spec: Dict[str, Any]) -> ComparisonReport: + """Compare two API specifications and detect changes.""" + self.old_spec = old_spec + self.new_spec = new_spec + self.report = ComparisonReport() + + # Compare different sections of the API specification + self._compare_info_section() + self._compare_servers_section() + self._compare_paths_section() + self._compare_components_section() + self._compare_security_section() + + # Calculate summary statistics + self.report.calculate_summary() + + return self.report + + def _compare_info_section(self) -> None: + """Compare API info sections.""" + old_info = self.old_spec.get('info', {}) + new_info = self.new_spec.get('info', {}) + + # Version comparison + old_version = old_info.get('version', '') + new_version = new_info.get('version', '') + + if old_version != new_version: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="versioning", + path="/info/version", + message=f"API version changed from '{old_version}' to '{new_version}'", + old_value=old_version, + new_value=new_version, + impact_description="Version change indicates API evolution" + )) + + # Title comparison + old_title = old_info.get('title', '') + new_title = new_info.get('title', '') + + if old_title != new_title: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="metadata", + path="/info/title", + message=f"API title changed from '{old_title}' to '{new_title}'", + old_value=old_title, + new_value=new_title, + impact_description="Title change is cosmetic and doesn't affect functionality" + )) + + def _compare_servers_section(self) -> None: + """Compare server configurations.""" + old_servers = self.old_spec.get('servers', []) + new_servers = self.new_spec.get('servers', []) + + old_urls = {server.get('url', '') for server in old_servers if isinstance(server, dict)} + new_urls = {server.get('url', '') for server in new_servers if isinstance(server, dict)} + + # Removed servers + removed_urls = old_urls - new_urls + for url in removed_urls: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="servers", + path="/servers", + message=f"Server URL removed: {url}", + old_value=url, + new_value=None, + migration_guide=f"Update client configurations to use alternative server URLs: {list(new_urls)}", + impact_description="Clients configured to use removed server URL will fail to connect" + )) + + # Added servers + added_urls = new_urls - old_urls + for url in added_urls: + self.report.add_change(Change( + change_type=ChangeType.ENHANCEMENT, + severity=ChangeSeverity.INFO, + category="servers", + path="/servers", + message=f"New server URL added: {url}", + old_value=None, + new_value=url, + impact_description="New server option provides additional deployment flexibility" + )) + + def _compare_paths_section(self) -> None: + """Compare API paths and operations.""" + old_paths = self.old_spec.get('paths', {}) + new_paths = self.new_spec.get('paths', {}) + + # Find removed, added, and modified paths + old_path_set = set(old_paths.keys()) + new_path_set = set(new_paths.keys()) + + removed_paths = old_path_set - new_path_set + added_paths = new_path_set - old_path_set + common_paths = old_path_set & new_path_set + + # Handle removed paths + for path in removed_paths: + old_operations = self._extract_operations(old_paths[path]) + for method in old_operations: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.CRITICAL, + category="endpoints", + path=f"/paths{path}", + message=f"Endpoint removed: {method.upper()} {path}", + old_value=f"{method.upper()} {path}", + new_value=None, + migration_guide=self._generate_endpoint_removal_migration(path, method, new_paths), + impact_description="Clients using this endpoint will receive 404 errors" + )) + + # Handle added paths + for path in added_paths: + new_operations = self._extract_operations(new_paths[path]) + for method in new_operations: + self.report.add_change(Change( + change_type=ChangeType.ENHANCEMENT, + severity=ChangeSeverity.INFO, + category="endpoints", + path=f"/paths{path}", + message=f"New endpoint added: {method.upper()} {path}", + old_value=None, + new_value=f"{method.upper()} {path}", + impact_description="New functionality available to clients" + )) + + # Handle modified paths + for path in common_paths: + self._compare_path_operations(path, old_paths[path], new_paths[path]) + + def _extract_operations(self, path_object: Dict[str, Any]) -> List[str]: + """Extract HTTP operations from a path object.""" + http_methods = {'get', 'post', 'put', 'patch', 'delete', 'head', 'options', 'trace'} + return [method for method in path_object.keys() if method.lower() in http_methods] + + def _compare_path_operations(self, path: str, old_path_obj: Dict, new_path_obj: Dict) -> None: + """Compare operations within a specific path.""" + old_operations = set(self._extract_operations(old_path_obj)) + new_operations = set(self._extract_operations(new_path_obj)) + + # Removed operations + removed_ops = old_operations - new_operations + for method in removed_ops: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.CRITICAL, + category="endpoints", + path=f"/paths{path}/{method}", + message=f"HTTP method removed: {method.upper()} {path}", + old_value=f"{method.upper()} {path}", + new_value=None, + migration_guide=self._generate_method_removal_migration(path, method, new_operations), + impact_description="Clients using this method will receive 405 Method Not Allowed errors" + )) + + # Added operations + added_ops = new_operations - old_operations + for method in added_ops: + self.report.add_change(Change( + change_type=ChangeType.ENHANCEMENT, + severity=ChangeSeverity.INFO, + category="endpoints", + path=f"/paths{path}/{method}", + message=f"New HTTP method added: {method.upper()} {path}", + old_value=None, + new_value=f"{method.upper()} {path}", + impact_description="New method provides additional functionality for this resource" + )) + + # Modified operations + common_ops = old_operations & new_operations + for method in common_ops: + self._compare_operation_details(path, method, old_path_obj[method], new_path_obj[method]) + + def _compare_operation_details(self, path: str, method: str, old_op: Dict, new_op: Dict) -> None: + """Compare details of individual operations.""" + operation_path = f"/paths{path}/{method}" + + # Compare parameters + self._compare_parameters(operation_path, old_op.get('parameters', []), new_op.get('parameters', [])) + + # Compare request body + self._compare_request_body(operation_path, old_op.get('requestBody'), new_op.get('requestBody')) + + # Compare responses + self._compare_responses(operation_path, old_op.get('responses', {}), new_op.get('responses', {})) + + # Compare security requirements + self._compare_security_requirements(operation_path, old_op.get('security'), new_op.get('security')) + + def _compare_parameters(self, base_path: str, old_params: List[Dict], new_params: List[Dict]) -> None: + """Compare operation parameters.""" + # Create lookup dictionaries + old_param_map = {(p.get('name'), p.get('in')): p for p in old_params} + new_param_map = {(p.get('name'), p.get('in')): p for p in new_params} + + old_param_keys = set(old_param_map.keys()) + new_param_keys = set(new_param_map.keys()) + + # Removed parameters + removed_params = old_param_keys - new_param_keys + for param_key in removed_params: + name, location = param_key + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="parameters", + path=f"{base_path}/parameters", + message=f"Parameter removed: {name} (in: {location})", + old_value=old_param_map[param_key], + new_value=None, + migration_guide=f"Remove '{name}' parameter from {location} when calling this endpoint", + impact_description="Clients sending this parameter may receive validation errors" + )) + + # Added parameters + added_params = new_param_keys - old_param_keys + for param_key in added_params: + name, location = param_key + new_param = new_param_map[param_key] + is_required = new_param.get('required', False) + + if is_required: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.CRITICAL, + category="parameters", + path=f"{base_path}/parameters", + message=f"New required parameter added: {name} (in: {location})", + old_value=None, + new_value=new_param, + migration_guide=f"Add required '{name}' parameter to {location} when calling this endpoint", + impact_description="Clients not providing this parameter will receive 400 Bad Request errors" + )) + else: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="parameters", + path=f"{base_path}/parameters", + message=f"New optional parameter added: {name} (in: {location})", + old_value=None, + new_value=new_param, + impact_description="Optional parameter provides additional functionality" + )) + + # Modified parameters + common_params = old_param_keys & new_param_keys + for param_key in common_params: + name, location = param_key + old_param = old_param_map[param_key] + new_param = new_param_map[param_key] + self._compare_parameter_details(base_path, name, location, old_param, new_param) + + def _compare_parameter_details(self, base_path: str, name: str, location: str, + old_param: Dict, new_param: Dict) -> None: + """Compare individual parameter details.""" + param_path = f"{base_path}/parameters/{name}" + + # Required status change + old_required = old_param.get('required', False) + new_required = new_param.get('required', False) + + if old_required != new_required: + if new_required: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="parameters", + path=param_path, + message=f"Parameter '{name}' is now required (was optional)", + old_value=old_required, + new_value=new_required, + migration_guide=f"Ensure '{name}' parameter is always provided when calling this endpoint", + impact_description="Clients not providing this parameter will receive validation errors" + )) + else: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="parameters", + path=param_path, + message=f"Parameter '{name}' is now optional (was required)", + old_value=old_required, + new_value=new_required, + impact_description="Parameter is now optional, providing more flexibility to clients" + )) + + # Schema/type changes + old_schema = old_param.get('schema', {}) + new_schema = new_param.get('schema', {}) + + if old_schema != new_schema: + self._compare_schemas(param_path, old_schema, new_schema, f"parameter '{name}'") + + def _compare_request_body(self, base_path: str, old_body: Optional[Dict], new_body: Optional[Dict]) -> None: + """Compare request body specifications.""" + body_path = f"{base_path}/requestBody" + + # Request body added + if old_body is None and new_body is not None: + is_required = new_body.get('required', False) + if is_required: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="request_body", + path=body_path, + message="Required request body added", + old_value=None, + new_value=new_body, + migration_guide="Include request body with appropriate content type when calling this endpoint", + impact_description="Clients not providing request body will receive validation errors" + )) + else: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="request_body", + path=body_path, + message="Optional request body added", + old_value=None, + new_value=new_body, + impact_description="Optional request body provides additional functionality" + )) + + # Request body removed + elif old_body is not None and new_body is None: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="request_body", + path=body_path, + message="Request body removed", + old_value=old_body, + new_value=None, + migration_guide="Remove request body when calling this endpoint", + impact_description="Clients sending request body may receive validation errors" + )) + + # Request body modified + elif old_body is not None and new_body is not None: + self._compare_request_body_details(body_path, old_body, new_body) + + def _compare_request_body_details(self, base_path: str, old_body: Dict, new_body: Dict) -> None: + """Compare request body details.""" + # Required status change + old_required = old_body.get('required', False) + new_required = new_body.get('required', False) + + if old_required != new_required: + if new_required: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="request_body", + path=base_path, + message="Request body is now required (was optional)", + old_value=old_required, + new_value=new_required, + migration_guide="Always include request body when calling this endpoint", + impact_description="Clients not providing request body will receive validation errors" + )) + else: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="request_body", + path=base_path, + message="Request body is now optional (was required)", + old_value=old_required, + new_value=new_required, + impact_description="Request body is now optional, providing more flexibility" + )) + + # Content type changes + old_content = old_body.get('content', {}) + new_content = new_body.get('content', {}) + self._compare_content_types(base_path, old_content, new_content, "request body") + + def _compare_responses(self, base_path: str, old_responses: Dict, new_responses: Dict) -> None: + """Compare response specifications.""" + responses_path = f"{base_path}/responses" + + old_status_codes = set(old_responses.keys()) + new_status_codes = set(new_responses.keys()) + + # Removed status codes + removed_codes = old_status_codes - new_status_codes + for code in removed_codes: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="responses", + path=f"{responses_path}/{code}", + message=f"Response status code {code} removed", + old_value=old_responses[code], + new_value=None, + migration_guide=f"Handle alternative status codes: {list(new_status_codes)}", + impact_description=f"Clients expecting status code {code} need to handle different responses" + )) + + # Added status codes + added_codes = new_status_codes - old_status_codes + for code in added_codes: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="responses", + path=f"{responses_path}/{code}", + message=f"New response status code {code} added", + old_value=None, + new_value=new_responses[code], + impact_description="New status code provides more specific response information" + )) + + # Modified responses + common_codes = old_status_codes & new_status_codes + for code in common_codes: + self._compare_response_details(responses_path, code, old_responses[code], new_responses[code]) + + def _compare_response_details(self, base_path: str, status_code: str, + old_response: Dict, new_response: Dict) -> None: + """Compare individual response details.""" + response_path = f"{base_path}/{status_code}" + + # Compare content types and schemas + old_content = old_response.get('content', {}) + new_content = new_response.get('content', {}) + + self._compare_content_types(response_path, old_content, new_content, f"response {status_code}") + + def _compare_content_types(self, base_path: str, old_content: Dict, new_content: Dict, context: str) -> None: + """Compare content types and their schemas.""" + old_types = set(old_content.keys()) + new_types = set(new_content.keys()) + + # Removed content types + removed_types = old_types - new_types + for content_type in removed_types: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="content_types", + path=f"{base_path}/content", + message=f"Content type '{content_type}' removed from {context}", + old_value=content_type, + new_value=None, + migration_guide=f"Use alternative content types: {list(new_types)}", + impact_description=f"Clients expecting '{content_type}' need to handle different formats" + )) + + # Added content types + added_types = new_types - old_types + for content_type in added_types: + self.report.add_change(Change( + change_type=ChangeType.ENHANCEMENT, + severity=ChangeSeverity.INFO, + category="content_types", + path=f"{base_path}/content", + message=f"New content type '{content_type}' added to {context}", + old_value=None, + new_value=content_type, + impact_description=f"Additional format option available for {context}" + )) + + # Modified schemas for common content types + common_types = old_types & new_types + for content_type in common_types: + old_media = old_content[content_type] + new_media = new_content[content_type] + + old_schema = old_media.get('schema', {}) + new_schema = new_media.get('schema', {}) + + if old_schema != new_schema: + schema_path = f"{base_path}/content/{content_type}/schema" + self._compare_schemas(schema_path, old_schema, new_schema, f"{context} ({content_type})") + + def _compare_schemas(self, base_path: str, old_schema: Dict, new_schema: Dict, context: str) -> None: + """Compare schema definitions.""" + # Type changes + old_type = old_schema.get('type') + new_type = new_schema.get('type') + + if old_type != new_type and old_type is not None and new_type is not None: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.CRITICAL, + category="schema", + path=base_path, + message=f"Schema type changed from '{old_type}' to '{new_type}' for {context}", + old_value=old_type, + new_value=new_type, + migration_guide=f"Update client code to handle {new_type} instead of {old_type}", + impact_description="Type change will break client parsing and validation" + )) + + # Property changes for object types + if old_schema.get('type') == 'object' and new_schema.get('type') == 'object': + self._compare_object_properties(base_path, old_schema, new_schema, context) + + # Array item changes + if old_schema.get('type') == 'array' and new_schema.get('type') == 'array': + old_items = old_schema.get('items', {}) + new_items = new_schema.get('items', {}) + if old_items != new_items: + self._compare_schemas(f"{base_path}/items", old_items, new_items, f"{context} items") + + def _compare_object_properties(self, base_path: str, old_schema: Dict, new_schema: Dict, context: str) -> None: + """Compare object schema properties.""" + old_props = old_schema.get('properties', {}) + new_props = new_schema.get('properties', {}) + old_required = set(old_schema.get('required', [])) + new_required = set(new_schema.get('required', [])) + + old_prop_names = set(old_props.keys()) + new_prop_names = set(new_props.keys()) + + # Removed properties + removed_props = old_prop_names - new_prop_names + for prop_name in removed_props: + severity = ChangeSeverity.CRITICAL if prop_name in old_required else ChangeSeverity.HIGH + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=severity, + category="schema", + path=f"{base_path}/properties", + message=f"Property '{prop_name}' removed from {context}", + old_value=old_props[prop_name], + new_value=None, + migration_guide=f"Remove references to '{prop_name}' property in client code", + impact_description="Clients expecting this property will receive incomplete data" + )) + + # Added properties + added_props = new_prop_names - old_prop_names + for prop_name in added_props: + if prop_name in new_required: + # This is handled separately in required field changes + pass + else: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="schema", + path=f"{base_path}/properties", + message=f"New optional property '{prop_name}' added to {context}", + old_value=None, + new_value=new_props[prop_name], + impact_description="New property provides additional data without breaking existing clients" + )) + + # Required field changes + added_required = new_required - old_required + removed_required = old_required - new_required + + for prop_name in added_required: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.CRITICAL, + category="schema", + path=f"{base_path}/properties", + message=f"Property '{prop_name}' is now required in {context}", + old_value=False, + new_value=True, + migration_guide=f"Ensure '{prop_name}' is always provided when sending {context}", + impact_description="Clients not providing this property will receive validation errors" + )) + + for prop_name in removed_required: + self.report.add_change(Change( + change_type=ChangeType.NON_BREAKING, + severity=ChangeSeverity.INFO, + category="schema", + path=f"{base_path}/properties", + message=f"Property '{prop_name}' is no longer required in {context}", + old_value=True, + new_value=False, + impact_description="Property is now optional, providing more flexibility" + )) + + # Modified properties + common_props = old_prop_names & new_prop_names + for prop_name in common_props: + old_prop = old_props[prop_name] + new_prop = new_props[prop_name] + if old_prop != new_prop: + self._compare_schemas(f"{base_path}/properties/{prop_name}", + old_prop, new_prop, f"{context}.{prop_name}") + + def _compare_security_requirements(self, base_path: str, old_security: Optional[List], + new_security: Optional[List]) -> None: + """Compare security requirements.""" + # Simplified security comparison - could be expanded + if old_security != new_security: + severity = ChangeSeverity.HIGH if new_security else ChangeSeverity.CRITICAL + change_type = ChangeType.BREAKING + + if old_security is None and new_security is not None: + message = "Security requirements added" + migration_guide = "Ensure proper authentication/authorization when calling this endpoint" + impact = "Endpoint now requires authentication" + elif old_security is not None and new_security is None: + message = "Security requirements removed" + migration_guide = "Authentication is no longer required for this endpoint" + impact = "Endpoint is now publicly accessible" + severity = ChangeSeverity.MEDIUM # Less severe, more permissive + else: + message = "Security requirements modified" + migration_guide = "Update authentication/authorization method for this endpoint" + impact = "Different authentication method required" + + self.report.add_change(Change( + change_type=change_type, + severity=severity, + category="security", + path=f"{base_path}/security", + message=message, + old_value=old_security, + new_value=new_security, + migration_guide=migration_guide, + impact_description=impact + )) + + def _compare_components_section(self) -> None: + """Compare components sections.""" + old_components = self.old_spec.get('components', {}) + new_components = self.new_spec.get('components', {}) + + # Compare schemas + old_schemas = old_components.get('schemas', {}) + new_schemas = new_components.get('schemas', {}) + + old_schema_names = set(old_schemas.keys()) + new_schema_names = set(new_schemas.keys()) + + # Removed schemas + removed_schemas = old_schema_names - new_schema_names + for schema_name in removed_schemas: + self.report.add_change(Change( + change_type=ChangeType.BREAKING, + severity=ChangeSeverity.HIGH, + category="components", + path=f"/components/schemas/{schema_name}", + message=f"Schema '{schema_name}' removed from components", + old_value=old_schemas[schema_name], + new_value=None, + migration_guide=f"Remove references to schema '{schema_name}' or use alternative schemas", + impact_description="References to this schema will fail validation" + )) + + # Added schemas + added_schemas = new_schema_names - old_schema_names + for schema_name in added_schemas: + self.report.add_change(Change( + change_type=ChangeType.ENHANCEMENT, + severity=ChangeSeverity.INFO, + category="components", + path=f"/components/schemas/{schema_name}", + message=f"New schema '{schema_name}' added to components", + old_value=None, + new_value=new_schemas[schema_name], + impact_description="New reusable schema available" + )) + + # Modified schemas + common_schemas = old_schema_names & new_schema_names + for schema_name in common_schemas: + old_schema = old_schemas[schema_name] + new_schema = new_schemas[schema_name] + if old_schema != new_schema: + self._compare_schemas(f"/components/schemas/{schema_name}", + old_schema, new_schema, f"schema '{schema_name}'") + + def _compare_security_section(self) -> None: + """Compare security definitions.""" + old_security_schemes = self.old_spec.get('components', {}).get('securitySchemes', {}) + new_security_schemes = self.new_spec.get('components', {}).get('securitySchemes', {}) + + if old_security_schemes != new_security_schemes: + # Simplified comparison - could be more detailed + self.report.add_change(Change( + change_type=ChangeType.POTENTIALLY_BREAKING, + severity=ChangeSeverity.MEDIUM, + category="security", + path="/components/securitySchemes", + message="Security scheme definitions changed", + old_value=old_security_schemes, + new_value=new_security_schemes, + migration_guide="Review authentication implementation for compatibility with new security schemes", + impact_description="Authentication mechanisms may have changed" + )) + + def _generate_endpoint_removal_migration(self, removed_path: str, method: str, + remaining_paths: Dict[str, Any]) -> str: + """Generate migration guide for removed endpoints.""" + # Look for similar endpoints + similar_paths = [] + path_segments = removed_path.strip('/').split('/') + + for existing_path in remaining_paths.keys(): + existing_segments = existing_path.strip('/').split('/') + if len(existing_segments) == len(path_segments): + # Check similarity + similarity = sum(1 for i, seg in enumerate(path_segments) + if i < len(existing_segments) and seg == existing_segments[i]) + if similarity >= len(path_segments) * 0.5: # At least 50% similar + similar_paths.append(existing_path) + + if similar_paths: + return f"Consider using alternative endpoints: {', '.join(similar_paths[:3])}" + else: + return "No direct replacement available. Review API documentation for alternative approaches." + + def _generate_method_removal_migration(self, path: str, removed_method: str, + remaining_methods: Set[str]) -> str: + """Generate migration guide for removed HTTP methods.""" + method_alternatives = { + 'get': ['head'], + 'post': ['put', 'patch'], + 'put': ['post', 'patch'], + 'patch': ['put', 'post'], + 'delete': [] + } + + alternatives = [] + for alt_method in method_alternatives.get(removed_method.lower(), []): + if alt_method in remaining_methods: + alternatives.append(alt_method.upper()) + + if alternatives: + return f"Use alternative methods: {', '.join(alternatives)}" + else: + return f"No alternative HTTP methods available for {path}" + + def generate_json_report(self) -> str: + """Generate JSON format report.""" + report_data = { + "summary": self.report.summary, + "hasBreakingChanges": self.report.has_breaking_changes(), + "changes": [change.to_dict() for change in self.report.changes] + } + + return json.dumps(report_data, indent=2) + + def generate_text_report(self) -> str: + """Generate human-readable text report.""" + lines = [ + "═══════════════════════════════════════════════════════════════", + " BREAKING CHANGE ANALYSIS REPORT", + "═══════════════════════════════════════════════════════════════", + "", + "SUMMARY:", + f" Total Changes: {self.report.summary.get('total_changes', 0)}", + f" 🔴 Breaking Changes: {self.report.summary.get('breaking_changes', 0)}", + f" 🟡 Potentially Breaking: {self.report.summary.get('potentially_breaking_changes', 0)}", + f" 🟢 Non-Breaking Changes: {self.report.summary.get('non_breaking_changes', 0)}", + f" ✨ Enhancements: {self.report.summary.get('enhancements', 0)}", + "", + "SEVERITY BREAKDOWN:", + f" 🚨 Critical: {self.report.summary.get('critical_severity', 0)}", + f" ⚠️ High: {self.report.summary.get('high_severity', 0)}", + f" ⚪ Medium: {self.report.summary.get('medium_severity', 0)}", + f" 🔵 Low: {self.report.summary.get('low_severity', 0)}", + f" ℹ️ Info: {self.report.summary.get('info_severity', 0)}", + "" + ] + + if not self.report.changes: + lines.extend([ + "🎉 No changes detected between the API versions!", + "" + ]) + else: + # Group changes by type and severity + breaking_changes = [c for c in self.report.changes if c.change_type == ChangeType.BREAKING] + potentially_breaking = [c for c in self.report.changes if c.change_type == ChangeType.POTENTIALLY_BREAKING] + non_breaking = [c for c in self.report.changes if c.change_type == ChangeType.NON_BREAKING] + enhancements = [c for c in self.report.changes if c.change_type == ChangeType.ENHANCEMENT] + + # Breaking changes section + if breaking_changes: + lines.extend([ + "🔴 BREAKING CHANGES:", + "═" * 60 + ]) + for change in sorted(breaking_changes, key=lambda x: x.severity.value): + self._add_change_to_report(lines, change) + lines.append("") + + # Potentially breaking changes section + if potentially_breaking: + lines.extend([ + "🟡 POTENTIALLY BREAKING CHANGES:", + "═" * 60 + ]) + for change in sorted(potentially_breaking, key=lambda x: x.severity.value): + self._add_change_to_report(lines, change) + lines.append("") + + # Non-breaking changes section + if non_breaking: + lines.extend([ + "🟢 NON-BREAKING CHANGES:", + "═" * 60 + ]) + for change in non_breaking: + self._add_change_to_report(lines, change) + lines.append("") + + # Enhancements section + if enhancements: + lines.extend([ + "✨ ENHANCEMENTS:", + "═" * 60 + ]) + for change in enhancements: + self._add_change_to_report(lines, change) + lines.append("") + + # Add overall assessment + lines.extend([ + "═══════════════════════════════════════════════════════════════", + "OVERALL ASSESSMENT:", + "═══════════════════════════════════════════════════════════════" + ]) + + if self.report.has_breaking_changes(): + breaking_count = self.report.summary.get('breaking_changes', 0) + potentially_breaking_count = self.report.summary.get('potentially_breaking_changes', 0) + + if breaking_count > 0: + lines.extend([ + f"⛔ MAJOR VERSION BUMP REQUIRED", + f" This API version contains {breaking_count} breaking changes that will", + f" definitely break existing clients. A major version bump is required.", + "" + ]) + elif potentially_breaking_count > 0: + lines.extend([ + f"⚠️ MINOR VERSION BUMP RECOMMENDED", + f" This API version contains {potentially_breaking_count} potentially breaking", + f" changes. Consider a minor version bump and communicate changes to clients.", + "" + ]) + else: + lines.extend([ + "✅ PATCH VERSION BUMP ACCEPTABLE", + " No breaking changes detected. This version is backward compatible", + " with existing clients.", + "" + ]) + + return "\n".join(lines) + + def _add_change_to_report(self, lines: List[str], change: Change) -> None: + """Add a change to the text report.""" + severity_icons = { + ChangeSeverity.CRITICAL: "🚨", + ChangeSeverity.HIGH: "⚠️ ", + ChangeSeverity.MEDIUM: "⚪", + ChangeSeverity.LOW: "🔵", + ChangeSeverity.INFO: "ℹ️ " + } + + icon = severity_icons.get(change.severity, "❓") + + lines.extend([ + f"{icon} {change.severity.value.upper()}: {change.message}", + f" Path: {change.path}", + f" Category: {change.category}" + ]) + + if change.impact_description: + lines.append(f" Impact: {change.impact_description}") + + if change.migration_guide: + lines.append(f" 💡 Migration: {change.migration_guide}") + + lines.append("") + + +def main(): + """Main CLI entry point.""" + parser = argparse.ArgumentParser( + description="Compare API specification versions to detect breaking changes", + formatter_class=argparse.RawDescriptionHelpFormatter, + epilog=""" +Examples: + python breaking_change_detector.py v1.json v2.json + python breaking_change_detector.py --format json v1.json v2.json > changes.json + python breaking_change_detector.py --output report.txt v1.json v2.json + """ + ) + + parser.add_argument( + 'old_spec', + help='Old API specification file (JSON format)' + ) + + parser.add_argument( + 'new_spec', + help='New API specification file (JSON format)' + ) + + parser.add_argument( + '--format', + choices=['text', 'json'], + default='text', + help='Output format (default: text)' + ) + + parser.add_argument( + '--output', + help='Output file (default: stdout)' + ) + + parser.add_argument( + '--exit-on-breaking', + action='store_true', + help='Exit with code 1 if breaking changes are detected' + ) + + args = parser.parse_args() + + # Load specification files + try: + with open(args.old_spec, 'r') as f: + old_spec = json.load(f) + except FileNotFoundError: + print(f"Error: Old specification file '{args.old_spec}' not found.", file=sys.stderr) + return 1 + except json.JSONDecodeError as e: + print(f"Error: Invalid JSON in '{args.old_spec}': {e}", file=sys.stderr) + return 1 + + try: + with open(args.new_spec, 'r') as f: + new_spec = json.load(f) + except FileNotFoundError: + print(f"Error: New specification file '{args.new_spec}' not found.", file=sys.stderr) + return 1 + except json.JSONDecodeError as e: + print(f"Error: Invalid JSON in '{args.new_spec}': {e}", file=sys.stderr) + return 1 + + # Initialize detector and compare specifications + detector = BreakingChangeDetector() + + try: + report = detector.compare_specs(old_spec, new_spec) + except Exception as e: + print(f"Error during comparison: {e}", file=sys.stderr) + return 1 + + # Generate report + if args.format == 'json': + output = detector.generate_json_report() + else: + output = detector.generate_text_report() + + # Write output + if args.output: + try: + with open(args.output, 'w') as f: + f.write(output) + print(f"Breaking change report written to {args.output}") + except IOError as e: + print(f"Error writing to '{args.output}': {e}", file=sys.stderr) + return 1 + else: + print(output) + + # Exit with appropriate code + if args.exit_on_breaking and report.has_breaking_changes(): + return 1 + + return 0 + + +if __name__ == '__main__': + sys.exit(main()) \ No newline at end of file diff --git a/skills/api-test-suite-builder/SKILL.md b/skills/api-test-suite-builder/SKILL.md new file mode 100644 index 00000000..4e4ce8a6 --- /dev/null +++ b/skills/api-test-suite-builder/SKILL.md @@ -0,0 +1,177 @@ +--- +name: "api-test-suite-builder" +description: "API Test Suite Builder" +--- + +# API Test Suite Builder + +**Tier:** POWERFUL +**Category:** Engineering +**Domain:** Testing / API Quality + +--- + +## Overview + +Scans API route definitions across frameworks (Next.js App Router, Express, FastAPI, Django REST) and +auto-generates comprehensive test suites covering auth, input validation, error codes, pagination, file +uploads, and rate limiting. Outputs ready-to-run test files for Vitest+Supertest (Node) or Pytest+httpx +(Python). + +--- + +## Core Capabilities + +- **Route detection** — scan source files to extract all API endpoints +- **Auth coverage** — valid/invalid/expired tokens, missing auth header +- **Input validation** — missing fields, wrong types, boundary values, injection attempts +- **Error code matrix** — 400/401/403/404/422/500 for each route +- **Pagination** — first/last/empty/oversized pages +- **File uploads** — valid, oversized, wrong MIME type, empty +- **Rate limiting** — burst detection, per-user vs global limits + +--- + +## When to Use + +- New API added — generate test scaffold before writing implementation (TDD) +- Legacy API with no tests — scan and generate baseline coverage +- API contract review — verify existing tests match current route definitions +- Pre-release regression check — ensure all routes have at least smoke tests +- Security audit prep — generate adversarial input tests + +--- + +## Route Detection + +### Next.js App Router +```bash +# Find all route handlers +find ./app/api -name "route.ts" -o -name "route.js" | sort + +# Extract HTTP methods from each route file +grep -rn "export async function\|export function" app/api/**/route.ts | \ + grep -oE "(GET|POST|PUT|PATCH|DELETE|HEAD|OPTIONS)" | sort -u + +# Full route map +find ./app/api -name "route.ts" | while read f; do + route=$(echo $f | sed 's|./app||' | sed 's|/route.ts||') + methods=$(grep -oE "export (async )?function (GET|POST|PUT|PATCH|DELETE)" "$f" | \ + grep -oE "(GET|POST|PUT|PATCH|DELETE)") + echo "$methods $route" +done +``` + +### Express +```bash +# Find all router files +find ./src -name "*.ts" -o -name "*.js" | xargs grep -l "router\.\(get\|post\|put\|delete\|patch\)" 2>/dev/null + +# Extract routes with line numbers +grep -rn "router\.\(get\|post\|put\|delete\|patch\)\|app\.\(get\|post\|put\|delete\|patch\)" \ + src/ --include="*.ts" | grep -oE "(get|post|put|delete|patch)\(['\"][^'\"]*['\"]" + +# Generate route map +grep -rn "router\.\|app\." src/ --include="*.ts" | \ + grep -oE "\.(get|post|put|delete|patch)\(['\"][^'\"]+['\"]" | \ + sed "s/\.\(.*\)('\(.*\)'/\U\1 \2/" +``` + +### FastAPI +```bash +# Find all route decorators +grep -rn "@app\.\|@router\." . --include="*.py" | \ + grep -E "@(app|router)\.(get|post|put|delete|patch)" + +# Extract with path and function name +grep -rn "@\(app\|router\)\.\(get\|post\|put\|delete\|patch\)" . --include="*.py" | \ + grep -oE "@(app|router)\.(get|post|put|delete|patch)\(['\"][^'\"]*['\"]" +``` + +### Django REST Framework +```bash +# urlpatterns extraction +grep -rn "path\|re_path\|url(" . --include="*.py" | grep "urlpatterns" -A 50 | \ + grep -E "path\(['\"]" | grep -oE "['\"][^'\"]+['\"]" | head -40 + +# ViewSet router registration +grep -rn "router\.register\|DefaultRouter\|SimpleRouter" . --include="*.py" +``` + +--- + +## Test Generation Patterns + +### Auth Test Matrix + +For every authenticated endpoint, generate: + +| Test Case | Expected Status | +|-----------|----------------| +| No Authorization header | 401 | +| Invalid token format | 401 | +| Valid token, wrong user role | 403 | +| Expired JWT token | 401 | +| Valid token, correct role | 2xx | +| Token from deleted user | 401 | + +### Input Validation Matrix + +For every POST/PUT/PATCH endpoint with a request body: + +| Test Case | Expected Status | +|-----------|----------------| +| Empty body `{}` | 400 or 422 | +| Missing required fields (one at a time) | 400 or 422 | +| Wrong type (string where int expected) | 400 or 422 | +| Boundary: value at min-1 | 400 or 422 | +| Boundary: value at min | 2xx | +| Boundary: value at max | 2xx | +| Boundary: value at max+1 | 400 or 422 | +| SQL injection in string field | 400 or 200 (sanitized) | +| XSS payload in string field | 400 or 200 (sanitized) | +| Null values for required fields | 400 or 422 | + +--- + +## Example Test Files +→ See references/example-test-files.md for details + +## Generating Tests from Route Scan + +When given a codebase, follow this process: + +1. **Scan routes** using the detection commands above +2. **Read each route handler** to understand: + - Expected request body schema + - Auth requirements (middleware, decorators) + - Return types and status codes + - Business rules (ownership, role checks) +3. **Generate test file** per route group using the patterns above +4. **Name tests descriptively**: `"returns 401 when token is expired"` not `"auth test 3"` +5. **Use factories/fixtures** for test data — never hardcode IDs +6. **Assert response shape**, not just status code + +--- + +## Common Pitfalls + +- **Testing only happy paths** — 80% of bugs live in error paths; test those first +- **Hardcoded test data IDs** — use factories/fixtures; IDs change between environments +- **Shared state between tests** — always clean up in afterEach/afterAll +- **Testing implementation, not behavior** — test what the API returns, not how it does it +- **Missing boundary tests** — off-by-one errors are extremely common in pagination and limits +- **Not testing token expiry** — expired tokens behave differently from invalid ones +- **Ignoring Content-Type** — test that API rejects wrong content types (xml when json expected) + +--- + +## Best Practices + +1. One describe block per endpoint — keeps failures isolated and readable +2. Seed minimal data — don't load the entire DB; create only what the test needs +3. Use `beforeAll` for shared setup, `afterAll` for cleanup — not `beforeEach` for expensive ops +4. Assert specific error messages/fields, not just status codes +5. Test that sensitive fields (password, secret) are never in responses +6. For auth tests, always test the "missing header" case separately from "invalid token" +7. Add rate limit tests last — they can interfere with other test suites if run in parallel diff --git a/skills/api-test-suite-builder/_meta.json b/skills/api-test-suite-builder/_meta.json new file mode 100644 index 00000000..cbbc0f2e --- /dev/null +++ b/skills/api-test-suite-builder/_meta.json @@ -0,0 +1,17 @@ +{ + "owner": "alirezarezvani", + "slug": "api-test-suite-builder", + "displayName": "api-test-suite-builder", + "latest": { + "version": "1.0.0", + "publishedAt": 1773242238273, + "commit": "https://github.com/openclaw/skills/commit/7844fd721019aa52c67fadd80046b21d63f1c543" + }, + "history": [ + { + "version": "2.1.1", + "publishedAt": 1773125604731, + "commit": "https://github.com/openclaw/skills/commit/693d0b2e323f0eb41cc8af963cbf77a23117cff8" + } + ] +} diff --git a/skills/api-test-suite-builder/references/example-test-files.md b/skills/api-test-suite-builder/references/example-test-files.md new file mode 100644 index 00000000..27b9d7e9 --- /dev/null +++ b/skills/api-test-suite-builder/references/example-test-files.md @@ -0,0 +1,508 @@ +# api-test-suite-builder reference + +## Example Test Files + +### Example 1 — Node.js: Vitest + Supertest (Next.js API Route) + +```typescript +// tests/api/users.test.ts +import { describe, it, expect, beforeAll, afterAll } from 'vitest' +import request from 'supertest' +import { createServer } from '@/test/helpers/server' +import { generateJWT, generateExpiredJWT } from '@/test/helpers/auth' +import { createTestUser, cleanupTestUsers } from '@/test/helpers/db' + +const app = createServer() + +describe('GET /api/users/:id', () => { + let validToken: string + let adminToken: string + let testUserId: string + + beforeAll(async () => { + const user = await createTestUser({ role: 'user' }) + const admin = await createTestUser({ role: 'admin' }) + testUserId = user.id + validToken = generateJWT(user) + adminToken = generateJWT(admin) + }) + + afterAll(async () => { + await cleanupTestUsers() + }) + + // --- Auth tests --- + it('returns 401 with no auth header', async () => { + const res = await request(app).get(`/api/users/${testUserId}`) + expect(res.status).toBe(401) + expect(res.body).toHaveProperty('error') + }) + + it('returns 401 with malformed token', async () => { + const res = await request(app) + .get(`/api/users/${testUserId}`) + .set('Authorization', 'Bearer not-a-real-jwt') + expect(res.status).toBe(401) + }) + + it('returns 401 with expired token', async () => { + const expiredToken = generateExpiredJWT({ id: testUserId }) + const res = await request(app) + .get(`/api/users/${testUserId}`) + .set('Authorization', `Bearer ${expiredToken}`) + expect(res.status).toBe(401) + expect(res.body.error).toMatch(/expired/i) + }) + + it('returns 403 when accessing another user\'s profile without admin', async () => { + const otherUser = await createTestUser({ role: 'user' }) + const otherToken = generateJWT(otherUser) + const res = await request(app) + .get(`/api/users/${testUserId}`) + .set('Authorization', `Bearer ${otherToken}`) + expect(res.status).toBe(403) + await cleanupTestUsers([otherUser.id]) + }) + + it('returns 200 with valid token for own profile', async () => { + const res = await request(app) + .get(`/api/users/${testUserId}`) + .set('Authorization', `Bearer ${validToken}`) + expect(res.status).toBe(200) + expect(res.body).toMatchObject({ id: testUserId }) + expect(res.body).not.toHaveProperty('password') + expect(res.body).not.toHaveProperty('hashedPassword') + }) + + it('returns 404 for non-existent user', async () => { + const res = await request(app) + .get('/api/users/00000000-0000-0000-0000-000000000000') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(404) + }) + + // --- Input validation --- + it('returns 400 for invalid UUID format', async () => { + const res = await request(app) + .get('/api/users/not-a-uuid') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(400) + }) +}) + +describe('POST /api/users', () => { + let adminToken: string + + beforeAll(async () => { + const admin = await createTestUser({ role: 'admin' }) + adminToken = generateJWT(admin) + }) + + afterAll(cleanupTestUsers) + + // --- Input validation --- + it('returns 422 when body is empty', async () => { + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({}) + expect(res.status).toBe(422) + expect(res.body.errors).toBeDefined() + }) + + it('returns 422 when email is missing', async () => { + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({ name: "test-user", role: 'user' }) + expect(res.status).toBe(422) + expect(res.body.errors).toContainEqual( + expect.objectContaining({ field: 'email' }) + ) + }) + + it('returns 422 for invalid email format', async () => { + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({ email: 'not-an-email', name: "test", role: 'user' }) + expect(res.status).toBe(422) + }) + + it('returns 422 for SQL injection attempt in email field', async () => { + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({ email: "' OR '1'='1", name: "hacker", role: 'user' }) + expect(res.status).toBe(422) + }) + + it('returns 409 when email already exists', async () => { + const existing = await createTestUser({ role: 'user' }) + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({ email: existing.email, name: "duplicate", role: 'user' }) + expect(res.status).toBe(409) + }) + + it('creates user successfully with valid data', async () => { + const res = await request(app) + .post('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + .send({ email: 'newuser@example.com', name: "new-user", role: 'user' }) + expect(res.status).toBe(201) + expect(res.body).toHaveProperty('id') + expect(res.body.email).toBe('newuser@example.com') + expect(res.body).not.toHaveProperty('password') + }) +}) + +describe('GET /api/users (pagination)', () => { + let adminToken: string + + beforeAll(async () => { + const admin = await createTestUser({ role: 'admin' }) + adminToken = generateJWT(admin) + // Create 15 test users for pagination + await Promise.all(Array.from({ length: 15 }, (_, i) => + createTestUser({ email: `pagtest${i}@example.com` }) + )) + }) + + afterAll(cleanupTestUsers) + + it('returns first page with default limit', async () => { + const res = await request(app) + .get('/api/users') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(200) + expect(res.body.data).toBeInstanceOf(Array) + expect(res.body).toHaveProperty('total') + expect(res.body).toHaveProperty('page') + expect(res.body).toHaveProperty('pageSize') + }) + + it('returns empty array for page beyond total', async () => { + const res = await request(app) + .get('/api/users?page=9999') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(200) + expect(res.body.data).toHaveLength(0) + }) + + it('returns 400 for negative page number', async () => { + const res = await request(app) + .get('/api/users?page=-1') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(400) + }) + + it('caps pageSize at maximum allowed value', async () => { + const res = await request(app) + .get('/api/users?pageSize=9999') + .set('Authorization', `Bearer ${adminToken}`) + expect(res.status).toBe(200) + expect(res.body.data.length).toBeLessThanOrEqual(100) + }) +}) +``` + +--- + +### Example 2 — Node.js: File Upload Tests + +```typescript +// tests/api/uploads.test.ts +import { describe, it, expect } from 'vitest' +import request from 'supertest' +import path from 'path' +import fs from 'fs' +import { createServer } from '@/test/helpers/server' +import { generateJWT } from '@/test/helpers/auth' +import { createTestUser } from '@/test/helpers/db' + +const app = createServer() + +describe('POST /api/upload', () => { + let validToken: string + + beforeAll(async () => { + const user = await createTestUser({ role: 'user' }) + validToken = generateJWT(user) + }) + + it('returns 401 without authentication', async () => { + const res = await request(app) + .post('/api/upload') + .attach('file', Buffer.from('test'), 'test.pdf') + expect(res.status).toBe(401) + }) + + it('returns 400 when no file attached', async () => { + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + expect(res.status).toBe(400) + expect(res.body.error).toMatch(/file/i) + }) + + it('returns 400 for unsupported file type (exe)', async () => { + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + .attach('file', Buffer.from('MZ fake exe'), { filename: "virusexe", contentType: 'application/octet-stream' }) + expect(res.status).toBe(400) + expect(res.body.error).toMatch(/type|format|allowed/i) + }) + + it('returns 413 for oversized file (>10MB)', async () => { + const largeBuf = Buffer.alloc(11 * 1024 * 1024) // 11MB + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + .attach('file', largeBuf, { filename: "largepdf", contentType: 'application/pdf' }) + expect(res.status).toBe(413) + }) + + it('returns 400 for empty file (0 bytes)', async () => { + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + .attach('file', Buffer.alloc(0), { filename: "emptypdf", contentType: 'application/pdf' }) + expect(res.status).toBe(400) + }) + + it('rejects MIME type spoofing (pdf extension but exe content)', async () => { + // Real malicious file: exe magic bytes but pdf extension + const fakeExe = Buffer.from('4D5A9000', 'hex') // MZ header + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + .attach('file', fakeExe, { filename: "documentpdf", contentType: 'application/pdf' }) + // Should detect magic bytes mismatch + expect([400, 415]).toContain(res.status) + }) + + it('accepts valid PDF file', async () => { + const pdfHeader = Buffer.from('%PDF-1.4 test content') + const res = await request(app) + .post('/api/upload') + .set('Authorization', `Bearer ${validToken}`) + .attach('file', pdfHeader, { filename: "validpdf", contentType: 'application/pdf' }) + expect(res.status).toBe(200) + expect(res.body).toHaveProperty('url') + expect(res.body).toHaveProperty('id') + }) +}) +``` + +--- + +### Example 3 — Python: Pytest + httpx (FastAPI) + +```python +# tests/api/test_items.py +import pytest +import httpx +from datetime import datetime, timedelta +import jwt + +BASE_URL = "http://localhost:8000" +JWT_SECRET = "test-secret" # use test config, never production secret + + +def make_token(user_id: str, role: str = "user", expired: bool = False) -> str: + exp = datetime.utcnow() + (timedelta(hours=-1) if expired else timedelta(hours=1)) + return jwt.encode( + {"sub": user_id, "role": role, "exp": exp}, + JWT_SECRET, + algorithm="HS256", + ) + + +@pytest.fixture +def client(): + with httpx.Client(base_url=BASE_URL) as c: + yield c + + +@pytest.fixture +def valid_token(): + return make_token("user-123", role="user") + + +@pytest.fixture +def admin_token(): + return make_token("admin-456", role="admin") + + +@pytest.fixture +def expired_token(): + return make_token("user-123", expired=True) + + +class TestGetItem: + def test_returns_401_without_auth(self, client): + res = client.get("/api/items/1") + assert res.status_code == 401 + + def test_returns_401_with_invalid_token(self, client): + res = client.get("/api/items/1", headers={"Authorization": "Bearer garbage"}) + assert res.status_code == 401 + + def test_returns_401_with_expired_token(self, client, expired_token): + res = client.get("/api/items/1", headers={"Authorization": f"Bearer {expired_token}"}) + assert res.status_code == 401 + assert "expired" in res.json().get("detail", "").lower() + + def test_returns_404_for_nonexistent_item(self, client, valid_token): + res = client.get( + "/api/items/99999999", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 404 + + def test_returns_400_for_invalid_id_format(self, client, valid_token): + res = client.get( + "/api/items/not-a-number", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code in (400, 422) + + def test_returns_200_with_valid_auth(self, client, valid_token, test_item): + res = client.get( + f"/api/items/{test_item['id']}", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 200 + data = res.json() + assert data["id"] == test_item["id"] + assert "password" not in data + + +class TestCreateItem: + def test_returns_422_with_empty_body(self, client, admin_token): + res = client.post( + "/api/items", + json={}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 422 + errors = res.json()["detail"] + assert len(errors) > 0 + + def test_returns_422_with_missing_required_field(self, client, admin_token): + res = client.post( + "/api/items", + json={"description": "no name field"}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 422 + fields = [e["loc"][-1] for e in res.json()["detail"]] + assert "name" in fields + + def test_returns_422_with_wrong_type(self, client, admin_token): + res = client.post( + "/api/items", + json={"name": "test", "price": "not-a-number"}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 422 + + @pytest.mark.parametrize("price", [-1, -0.01]) + def test_returns_422_for_negative_price(self, client, admin_token, price): + res = client.post( + "/api/items", + json={"name": "test", "price": price}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 422 + + def test_returns_422_for_price_exceeding_max(self, client, admin_token): + res = client.post( + "/api/items", + json={"name": "test", "price": 1_000_001}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 422 + + def test_creates_item_successfully(self, client, admin_token): + res = client.post( + "/api/items", + json={"name": "New Widget", "price": 9.99, "category": "tools"}, + headers={"Authorization": f"Bearer {admin_token}"}, + ) + assert res.status_code == 201 + data = res.json() + assert "id" in data + assert data["name"] == "New Widget" + + def test_returns_403_for_non_admin(self, client, valid_token): + res = client.post( + "/api/items", + json={"name": "test", "price": 1.0}, + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 403 + + +class TestPagination: + def test_returns_paginated_response(self, client, valid_token): + res = client.get( + "/api/items?page=1&size=10", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 200 + data = res.json() + assert "items" in data + assert "total" in data + assert "page" in data + assert len(data["items"]) <= 10 + + def test_empty_result_for_out_of_range_page(self, client, valid_token): + res = client.get( + "/api/items?page=99999", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 200 + assert res.json()["items"] == [] + + def test_returns_422_for_page_zero(self, client, valid_token): + res = client.get( + "/api/items?page=0", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 422 + + def test_caps_page_size_at_maximum(self, client, valid_token): + res = client.get( + "/api/items?size=9999", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + assert res.status_code == 200 + assert len(res.json()["items"]) <= 100 # max page size + + +class TestRateLimiting: + def test_rate_limit_after_burst(self, client, valid_token): + responses = [] + for _ in range(60): # exceed typical 50/min limit + res = client.get( + "/api/items", + headers={"Authorization": f"Bearer {valid_token}"}, + ) + responses.append(res.status_code) + if res.status_code == 429: + break + assert 429 in responses, "Rate limit was not triggered" + + def test_rate_limit_response_has_retry_after(self, client, valid_token): + for _ in range(60): + res = client.get("/api/items", headers={"Authorization": f"Bearer {valid_token}"}) + if res.status_code == 429: + assert "Retry-After" in res.headers or "retry_after" in res.json() + break +``` + +--- diff --git a/skills/app-store-optimization/HOW_TO_USE.md b/skills/app-store-optimization/HOW_TO_USE.md new file mode 100644 index 00000000..67e68a8d --- /dev/null +++ b/skills/app-store-optimization/HOW_TO_USE.md @@ -0,0 +1,281 @@ +# How to Use the App Store Optimization Skill + +Hey Claude—I just added the "app-store-optimization" skill. Can you help me optimize my app's presence on the App Store and Google Play? + +## Example Invocations + +### Keyword Research + +**Example 1: Basic Keyword Research** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you research the best keywords for my productivity app? I'm targeting professionals who need task management and team collaboration features. +``` + +**Example 2: Competitive Keyword Analysis** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you analyze keywords that Todoist, Asana, and Monday.com are using? I want to find gaps and opportunities for my project management app. +``` + +### Metadata Optimization + +**Example 3: Optimize App Title** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you optimize my app title for the Apple App Store? My app is called "TaskFlow" and I want to rank for "task manager", "productivity", and "team collaboration". The title needs to be under 30 characters. +``` + +**Example 4: Full Metadata Package** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you create optimized metadata for both Apple App Store and Google Play Store? Here's my app info: +- Name: TaskFlow +- Category: Productivity +- Key features: AI task prioritization, team collaboration, calendar integration +- Target keywords: task manager, productivity app, team tasks +``` + +### Competitor Analysis + +**Example 5: Analyze Top Competitors** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you analyze the ASO strategies of the top 5 productivity apps in the App Store? I want to understand their title strategies, keyword usage, and visual asset approaches. +``` + +**Example 6: Identify Competitive Gaps** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you compare my app's ASO performance against competitors and identify what I'm missing? Here's my current metadata: [paste metadata] +``` + +### ASO Score Calculation + +**Example 7: Calculate Overall ASO Health** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you calculate my app's ASO health score? Here are my metrics: +- Average rating: 4.2 stars +- Total ratings: 3,500 +- Keywords in top 10: 3 +- Keywords in top 50: 12 +- Conversion rate: 4.5% +``` + +**Example 8: Identify Improvement Areas** +``` +Hey Claude—I just added the "app-store-optimization" skill. My ASO score is 62/100. Can you tell me which areas I should focus on first to improve my rankings and downloads? +``` + +### A/B Testing + +**Example 9: Plan Icon Test** +``` +Hey Claude—I just added the "app-store-optimization" skill. I want to A/B test two different app icons. My current conversion rate is 5%. Can you help me plan the test, calculate required sample size, and determine how long to run it? +``` + +**Example 10: Analyze Test Results** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you analyze my A/B test results? +- Variant A (control): 2,500 visitors, 125 installs +- Variant B (new icon): 2,500 visitors, 150 installs +Is this statistically significant? Should I implement variant B? +``` + +### Localization + +**Example 11: Plan Localization Strategy** +``` +Hey Claude—I just added the "app-store-optimization" skill. I currently only have English metadata. Which markets should I localize for first? I'm a bootstrapped startup with moderate budget. +``` + +**Example 12: Translate Metadata** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you help me translate my app metadata to Spanish for the Mexico market? Here's my English metadata: [paste metadata]. Check if it fits within character limits. +``` + +### Review Analysis + +**Example 13: Analyze User Reviews** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you analyze my recent reviews and tell me: +- Overall sentiment (positive/negative ratio) +- Most common complaints +- Most requested features +- Bugs that need immediate fixing +``` + +**Example 14: Generate Review Response Templates** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you create professional response templates for: +- Users reporting crashes +- Feature requests +- Positive 5-star reviews +- General complaints +``` + +### Launch Planning + +**Example 15: Pre-Launch Checklist** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you generate a comprehensive pre-launch checklist for both Apple App Store and Google Play Store? My launch date is December 1, 2025. +``` + +**Example 16: Optimize Launch Timing** +``` +Hey Claude—I just added the "app-store-optimization" skill. What's the best day and time to launch my fitness app? I want to maximize visibility and downloads in the first week. +``` + +**Example 17: Plan Seasonal Campaign** +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you identify seasonal opportunities for my fitness app? It's currently October—what campaigns should I run for the next 6 months? +``` + +## What to Provide + +### For Keyword Research +- App name and category +- Target audience description +- Key features and unique value proposition +- Competitor apps (optional) +- Geographic markets to target + +### For Metadata Optimization +- Current app name +- Platform (Apple, Google, or both) +- Target keywords (prioritized list) +- Key features and benefits +- Target audience +- Current metadata (for optimization) + +### For Competitor Analysis +- Your app category +- List of competitor app names or IDs +- Platform (Apple or Google) +- Specific aspects to analyze (keywords, visuals, ratings) + +### For ASO Score Calculation +- Metadata quality metrics (title length, description length, keyword density) +- Rating data (average rating, total ratings, recent ratings) +- Keyword rankings (top 10, top 50, top 100 counts) +- Conversion metrics (impression-to-install rate, downloads) + +### For A/B Testing +- Test type (icon, screenshot, title, description) +- Control variant details +- Test variant details +- Baseline conversion rate +- For results analysis: visitor and conversion counts for both variants + +### For Localization +- Current market and language +- Budget level (low, medium, high) +- Target number of markets +- Current metadata text for translation + +### For Review Analysis +- Recent reviews (text, rating, date) +- Platform (Apple or Google) +- Time period to analyze +- Specific focus (bugs, features, sentiment) + +### For Launch Planning +- Platform (Apple, Google, or both) +- Target launch date +- App category +- App information (name, features, target audience) + +## What You'll Get + +### Keyword Research Output +- Prioritized keyword list with search volume estimates +- Competition level analysis +- Relevance scores +- Long-tail keyword opportunities +- Strategic recommendations + +### Metadata Optimization Output +- Optimized titles (multiple options) +- Optimized descriptions (short and full) +- Keyword field optimization (Apple) +- Character count validation +- Keyword density analysis +- Before/after comparison + +### Competitor Analysis Output +- Ranked competitors by ASO strength +- Common keyword patterns +- Keyword gaps and opportunities +- Visual asset assessment +- Best practices identified +- Actionable recommendations + +### ASO Score Output +- Overall score (0-100) +- Breakdown by category (metadata, ratings, keywords, conversion) +- Strengths and weaknesses +- Prioritized action items +- Expected impact of improvements + +### A/B Test Output +- Test design with hypothesis +- Required sample size calculation +- Duration estimates +- Statistical significance analysis +- Implementation recommendations +- Learnings and insights + +### Localization Output +- Prioritized target markets +- Estimated translation costs +- ROI projections +- Character limit validation for each language +- Cultural adaptation recommendations +- Phased implementation plan + +### Review Analysis Output +- Sentiment distribution (positive/neutral/negative) +- Common themes and topics +- Top issues requiring fixes +- Most requested features +- Response templates +- Trend analysis over time + +### Launch Planning Output +- Platform-specific checklists (Apple, Google, Universal) +- Timeline with milestones +- Compliance validation +- Optimal launch timing recommendations +- Seasonal campaign opportunities +- Update cadence planning + +## Tips for Best Results + +1. **Be Specific**: Provide as much detail about your app as possible +2. **Include Context**: Share your goals (increase downloads, improve ranking, boost conversion) +3. **Provide Data**: Real metrics enable more accurate analysis +4. **Iterate**: Start with keyword research, then optimize metadata, then test +5. **Track Results**: Monitor changes after implementing recommendations +6. **Stay Compliant**: Always verify recommendations against current App Store/Play Store guidelines +7. **Test First**: Use A/B testing before making major metadata changes +8. **Localize Strategically**: Start with highest-ROI markets first +9. **Respond to Reviews**: Use provided templates to engage with users +10. **Plan Ahead**: Use launch checklists and timelines to avoid last-minute rushes + +## Common Workflows + +### New App Launch +1. Keyword research → Competitor analysis → Metadata optimization → Pre-launch checklist → Launch timing optimization + +### Improving Existing App +1. ASO score calculation → Identify gaps → Metadata optimization → A/B testing → Review analysis → Implement changes + +### International Expansion +1. Localization planning → Market prioritization → Metadata translation → ROI analysis → Phased rollout + +### Ongoing Optimization +1. Monthly keyword ranking tracking → Quarterly metadata updates → Continuous A/B testing → Review monitoring → Seasonal campaigns + +## Need Help? + +If you need clarification on any aspect of ASO or want to combine multiple analyses, just ask! For example: + +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you create a complete ASO strategy for my new productivity app? I need keyword research, optimized metadata for both stores, a pre-launch checklist, and launch timing recommendations. +``` + +The skill can handle comprehensive, multi-phase ASO projects as well as specific tactical optimizations. diff --git a/skills/app-store-optimization/README.md b/skills/app-store-optimization/README.md new file mode 100644 index 00000000..d22441d2 --- /dev/null +++ b/skills/app-store-optimization/README.md @@ -0,0 +1,430 @@ +# App Store Optimization (ASO) Skill + +**Version**: 1.0.0 +**Last Updated**: November 7, 2025 +**Author**: Claude Skills Factory + +## Overview + +A comprehensive App Store Optimization (ASO) skill that provides complete capabilities for researching, optimizing, and tracking mobile app performance on the Apple App Store and Google Play Store. This skill empowers app developers and marketers to maximize their app's visibility, downloads, and success in competitive app marketplaces. + +## What This Skill Does + +This skill provides end-to-end ASO capabilities across seven key areas: + +1. **Research & Analysis**: Keyword research, competitor analysis, market trends, review sentiment +2. **Metadata Optimization**: Title, description, keywords with platform-specific character limits +3. **Conversion Optimization**: A/B testing framework, visual asset optimization +4. **Rating & Review Management**: Sentiment analysis, response strategies, issue identification +5. **Launch & Update Strategies**: Pre-launch checklists, timing optimization, update planning +6. **Analytics & Tracking**: ASO scoring, keyword rankings, performance benchmarking +7. **Localization**: Multi-language strategy, translation management, ROI analysis + +## Key Features + +### Comprehensive Keyword Research +- Search volume and competition analysis +- Long-tail keyword discovery +- Competitor keyword extraction +- Keyword difficulty scoring +- Strategic prioritization + +### Platform-Specific Metadata Optimization +- **Apple App Store**: + - Title (30 chars) + - Subtitle (30 chars) + - Promotional Text (170 chars) + - Description (4000 chars) + - Keywords field (100 chars) +- **Google Play Store**: + - Title (50 chars) + - Short Description (80 chars) + - Full Description (4000 chars) +- Character limit validation +- Keyword density analysis +- Multiple optimization strategies + +### Competitor Intelligence +- Automated competitor discovery +- Metadata strategy analysis +- Visual asset assessment +- Gap identification +- Competitive positioning + +### ASO Health Scoring +- 0-100 overall score +- Four-category breakdown (Metadata, Ratings, Keywords, Conversion) +- Strengths and weaknesses identification +- Prioritized action recommendations +- Expected impact estimates + +### Scientific A/B Testing +- Test design and hypothesis formulation +- Sample size calculation +- Statistical significance analysis +- Duration estimation +- Implementation recommendations + +### Global Localization +- Market prioritization (Tier 1/2/3) +- Translation cost estimation +- Character limit adaptation by language +- Cultural keyword considerations +- ROI analysis + +### Review Intelligence +- Sentiment analysis +- Common theme extraction +- Bug and issue identification +- Feature request clustering +- Professional response templates + +### Launch Planning +- Platform-specific checklists +- Timeline generation +- Compliance validation +- Optimal timing recommendations +- Seasonal campaign planning + +## Python Modules + +This skill includes 8 powerful Python modules: + +### 1. keyword_analyzer.py +**Purpose**: Analyzes keywords for search volume, competition, and relevance + +**Key Functions**: +- `analyze_keyword()`: Single keyword analysis +- `compare_keywords()`: Multi-keyword comparison and ranking +- `find_long_tail_opportunities()`: Generate long-tail variations +- `calculate_keyword_density()`: Analyze keyword usage in text +- `extract_keywords_from_text()`: Extract keywords from reviews/descriptions + +### 2. metadata_optimizer.py +**Purpose**: Optimizes titles, descriptions, keywords with character limit validation + +**Key Functions**: +- `optimize_title()`: Generate optimal title options +- `optimize_description()`: Create conversion-focused descriptions +- `optimize_keyword_field()`: Maximize Apple's 100-char keyword field +- `validate_character_limits()`: Ensure platform compliance +- `calculate_keyword_density()`: Analyze keyword integration + +### 3. competitor_analyzer.py +**Purpose**: Analyzes competitor ASO strategies + +**Key Functions**: +- `analyze_competitor()`: Single competitor deep-dive +- `compare_competitors()`: Multi-competitor analysis +- `identify_gaps()`: Find competitive opportunities +- `_calculate_competitive_strength()`: Score competitor ASO quality + +### 4. aso_scorer.py +**Purpose**: Calculates comprehensive ASO health score + +**Key Functions**: +- `calculate_overall_score()`: 0-100 ASO health score +- `score_metadata_quality()`: Evaluate metadata optimization +- `score_ratings_reviews()`: Assess rating quality and volume +- `score_keyword_performance()`: Analyze ranking positions +- `score_conversion_metrics()`: Evaluate conversion rates +- `generate_recommendations()`: Prioritized improvement actions + +### 5. ab_test_planner.py +**Purpose**: Plans and tracks A/B tests for ASO elements + +**Key Functions**: +- `design_test()`: Create test hypothesis and structure +- `calculate_sample_size()`: Determine required visitors +- `calculate_significance()`: Assess statistical validity +- `track_test_results()`: Monitor ongoing tests +- `generate_test_report()`: Create comprehensive test reports + +### 6. localization_helper.py +**Purpose**: Manages multi-language ASO optimization + +**Key Functions**: +- `identify_target_markets()`: Prioritize localization markets +- `translate_metadata()`: Adapt metadata for languages +- `adapt_keywords()`: Cultural keyword adaptation +- `validate_translations()`: Character limit validation +- `calculate_localization_roi()`: Estimate investment returns + +### 7. review_analyzer.py +**Purpose**: Analyzes user reviews for actionable insights + +**Key Functions**: +- `analyze_sentiment()`: Calculate sentiment distribution +- `extract_common_themes()`: Identify frequent topics +- `identify_issues()`: Surface bugs and problems +- `find_feature_requests()`: Extract desired features +- `track_sentiment_trends()`: Monitor changes over time +- `generate_response_templates()`: Create review responses + +### 8. launch_checklist.py +**Purpose**: Generates comprehensive launch and update checklists + +**Key Functions**: +- `generate_prelaunch_checklist()`: Complete submission validation +- `validate_app_store_compliance()`: Check guidelines compliance +- `create_update_plan()`: Plan update cadence +- `optimize_launch_timing()`: Recommend launch dates +- `plan_seasonal_campaigns()`: Identify seasonal opportunities + +## Installation + +### For Claude Code (Desktop/CLI) + +#### Project-Level Installation +```bash +# Copy skill folder to project +cp -r app-store-optimization /path/to/your/project/.claude/skills/ + +# Claude will auto-load the skill when working in this project +``` + +#### User-Level Installation (Available in All Projects) +```bash +# Copy skill folder to user-level skills +cp -r app-store-optimization ~/.claude/skills/ + +# Claude will load this skill in all your projects +``` + +### For Claude Apps (Browser) + +1. Use the `skill-creator` skill to import the skill +2. Or manually import via Claude Apps interface + +### Verification + +To verify installation: +```bash +# Check if skill folder exists +ls ~/.claude/skills/app-store-optimization/ + +# You should see: +# SKILL.md +# keyword_analyzer.py +# metadata_optimizer.py +# competitor_analyzer.py +# aso_scorer.py +# ab_test_planner.py +# localization_helper.py +# review_analyzer.py +# launch_checklist.py +# sample_input.json +# expected_output.json +# HOW_TO_USE.md +# README.md +``` + +## Usage Examples + +### Example 1: Complete Keyword Research + +``` +Hey Claude—I just added the "app-store-optimization" skill. Can you research keywords for my fitness app? I'm targeting people who want home workouts, yoga, and meal planning. Analyze top competitors like Nike Training Club and Peloton. +``` + +**What Claude will do**: +- Use `keyword_analyzer.py` to research keywords +- Use `competitor_analyzer.py` to analyze Nike Training Club and Peloton +- Provide prioritized keyword list with search volumes, competition levels +- Identify gaps and long-tail opportunities +- Recommend primary keywords for title and secondary keywords for description + +### Example 2: Optimize App Store Metadata + +``` +Hey Claude—I just added the "app-store-optimization" skill. Optimize my app's metadata for both Apple App Store and Google Play Store: +- App: FitFlow +- Category: Health & Fitness +- Features: AI workout plans, nutrition tracking, progress photos +- Keywords: fitness app, workout planner, home fitness +``` + +**What Claude will do**: +- Use `metadata_optimizer.py` to create optimized titles (multiple options) +- Generate platform-specific descriptions (short and full) +- Optimize Apple's 100-character keyword field +- Validate all character limits +- Calculate keyword density +- Provide before/after comparison + +### Example 3: Calculate ASO Health Score + +``` +Hey Claude—I just added the "app-store-optimization" skill. Calculate my app's ASO score: +- Average rating: 4.3 stars (8,200 ratings) +- Keywords in top 10: 4 +- Keywords in top 50: 15 +- Conversion rate: 3.8% +- Title: "FitFlow - Home Workouts" +- Description: 1,500 characters with 3 keyword mentions +``` + +**What Claude will do**: +- Use `aso_scorer.py` to calculate overall score (0-100) +- Break down by category (Metadata: X/25, Ratings: X/25, Keywords: X/25, Conversion: X/25) +- Identify strengths and weaknesses +- Generate prioritized recommendations +- Estimate impact of improvements + +### Example 4: A/B Test Planning + +``` +Hey Claude—I just added the "app-store-optimization" skill. I want to A/B test my app icon. My current conversion rate is 4.2%. How many visitors do I need and how long should I run the test? +``` + +**What Claude will do**: +- Use `ab_test_planner.py` to design test +- Calculate required sample size (based on minimum detectable effect) +- Estimate test duration for low/medium/high traffic scenarios +- Provide test structure and success metrics +- Explain how to analyze results + +### Example 5: Review Sentiment Analysis + +``` +Hey Claude—I just added the "app-store-optimization" skill. Analyze my last 500 reviews and tell me: +- Overall sentiment +- Most common complaints +- Top feature requests +- Bugs needing immediate fixes +``` + +**What Claude will do**: +- Use `review_analyzer.py` to process reviews +- Calculate sentiment distribution +- Extract common themes +- Identify and prioritize issues +- Cluster feature requests +- Generate response templates + +### Example 6: Pre-Launch Checklist + +``` +Hey Claude—I just added the "app-store-optimization" skill. Generate a complete pre-launch checklist for both app stores. My launch date is March 15, 2026. +``` + +**What Claude will do**: +- Use `launch_checklist.py` to generate checklists +- Create Apple App Store checklist (metadata, assets, technical, legal) +- Create Google Play Store checklist (metadata, assets, technical, legal) +- Add universal checklist (marketing, QA, support) +- Generate timeline with milestones +- Calculate completion percentage + +## Best Practices + +### Keyword Research +1. Start with 20-30 seed keywords +2. Analyze top 5 competitors in your category +3. Balance high-volume and long-tail keywords +4. Prioritize relevance over search volume +5. Update keyword research quarterly + +### Metadata Optimization +1. Front-load keywords in title (first 15 characters most important) +2. Use every available character (don't waste space) +3. Write for humans first, search engines second +4. A/B test major changes before committing +5. Update descriptions with each major release + +### A/B Testing +1. Test one element at a time (icon vs. screenshots vs. title) +2. Run tests to statistical significance (90%+ confidence) +3. Test high-impact elements first (icon has biggest impact) +4. Allow sufficient duration (at least 1 week, preferably 2-3) +5. Document learnings for future tests + +### Localization +1. Start with top 5 revenue markets (US, China, Japan, Germany, UK) +2. Use professional translators, not machine translation +3. Test translations with native speakers +4. Adapt keywords for cultural context +5. Monitor ROI by market + +### Review Management +1. Respond to reviews within 24-48 hours +2. Always be professional, even with negative reviews +3. Address specific issues raised +4. Thank users for positive feedback +5. Use insights to prioritize product improvements + +## Technical Requirements + +- **Python**: 3.7+ (for Python modules) +- **Platform Support**: Apple App Store, Google Play Store +- **Data Formats**: JSON input/output +- **Dependencies**: Standard library only (no external packages required) + +## Limitations + +### Data Dependencies +- Keyword search volumes are estimates (no official Apple/Google data) +- Competitor data limited to publicly available information +- Review analysis requires access to public reviews +- Historical data may not be available for new apps + +### Platform Constraints +- Apple: Metadata changes require app submission (except Promotional Text) +- Google: Metadata changes take 1-2 hours to index +- A/B testing requires significant traffic for statistical significance +- Store algorithms are proprietary and change without notice + +### Scope +- Does not include paid user acquisition (Apple Search Ads, Google Ads) +- Does not cover in-app analytics implementation +- Does not handle technical app development +- Focuses on organic discovery and conversion optimization + +## Troubleshooting + +### Issue: Python modules not found +**Solution**: Ensure all .py files are in the same directory as SKILL.md + +### Issue: Character limit validation failing +**Solution**: Check that you're using the correct platform ('apple' or 'google') + +### Issue: Keyword research returning limited results +**Solution**: Provide more context about your app, features, and target audience + +### Issue: ASO score seems inaccurate +**Solution**: Ensure you're providing accurate metrics (ratings, keyword rankings, conversion rate) + +## Version History + +### Version 1.0.0 (November 7, 2025) +- Initial release +- 8 Python modules with comprehensive ASO capabilities +- Support for both Apple App Store and Google Play Store +- Keyword research, metadata optimization, competitor analysis +- ASO scoring, A/B testing, localization, review analysis +- Launch planning and seasonal campaign tools + +## Support & Feedback + +This skill is designed to help app developers and marketers succeed in competitive app marketplaces. For the best results: + +1. Provide detailed context about your app +2. Include specific metrics when available +3. Ask follow-up questions for clarification +4. Iterate based on results + +## Credits + +Developed by Claude Skills Factory +Based on industry-standard ASO best practices +Platform requirements current as of November 2025 + +## License + +This skill is provided as-is for use with Claude Code and Claude Apps. Customize and extend as needed for your specific use cases. + +--- + +**Ready to optimize your app?** Start with keyword research, then move to metadata optimization, and finally implement A/B testing for continuous improvement. The skill handles everything from pre-launch planning to ongoing optimization. + +For detailed usage examples, see [HOW_TO_USE.md](HOW_TO_USE.md). diff --git a/skills/app-store-optimization/SKILL.md b/skills/app-store-optimization/SKILL.md new file mode 100644 index 00000000..2b07daf5 --- /dev/null +++ b/skills/app-store-optimization/SKILL.md @@ -0,0 +1,486 @@ +--- +name: "app-store-optimization" +description: App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklists, and tracking ranking changes. +triggers: + - ASO + - app store optimization + - app store ranking + - app keywords + - app metadata + - play store optimization + - app store listing + - improve app rankings + - app visibility + - app store SEO + - mobile app marketing + - app conversion rate +--- + +# App Store Optimization (ASO) + +--- + +## Keyword Research Workflow + +Discover and evaluate keywords that drive app store visibility. + +### Workflow: Conduct Keyword Research + +1. Define target audience and core app functions: + - Primary use case (what problem does the app solve) + - Target user demographics + - Competitive category +2. Generate seed keywords from: + - App features and benefits + - User language (not developer terminology) + - App store autocomplete suggestions +3. Expand keyword list using: + - Modifiers (free, best, simple) + - Actions (create, track, organize) + - Audiences (for students, for teams, for business) +4. Evaluate each keyword: + - Search volume (estimated monthly searches) + - Competition (number and quality of ranking apps) + - Relevance (alignment with app function) +5. Score and prioritize keywords: + - Primary: Title and keyword field (iOS) + - Secondary: Subtitle and short description + - Tertiary: Full description only +6. Map keywords to metadata locations +7. Document keyword strategy for tracking +8. **Validation:** Keywords scored; placement mapped; no competitor brand names included; no plurals in iOS keyword field + +### Keyword Evaluation Criteria + +| Factor | Weight | High Score Indicators | +|--------|--------|----------------------| +| Relevance | 35% | Describes core app function | +| Volume | 25% | 10,000+ monthly searches | +| Competition | 25% | Top 10 apps have <4.5 avg rating | +| Conversion | 15% | Transactional intent ("best X app") | + +### Keyword Placement Priority + +| Location | Search Weight | +|----------|---------------| +| App Title | Highest | +| Subtitle (iOS) | High | +| Keyword Field (iOS) | High | +| Short Description (Android) | High | +| Full Description | Medium | + +See: [references/keyword-research-guide.md](references/keyword-research-guide.md) + +--- + +## Metadata Optimization Workflow + +Optimize app store listing elements for search ranking and conversion. + +### Workflow: Optimize App Metadata + +1. Audit current metadata against platform limits: + - Title character count and keyword presence + - Subtitle/short description usage + - Keyword field efficiency (iOS) + - Description keyword density +2. Optimize title following formula: + ``` + [Brand Name] - [Primary Keyword] [Secondary Keyword] + ``` +3. Write subtitle (iOS) or short description (Android): + - Focus on primary benefit + - Include secondary keyword + - Use action verbs +4. Optimize keyword field (iOS only): + - Remove duplicates from title + - Remove plurals (Apple indexes both forms) + - No spaces after commas + - Prioritize by score +5. Rewrite full description: + - Hook paragraph with value proposition + - Feature bullets with keywords + - Social proof section + - Call to action +6. Validate character counts for each field +7. Calculate keyword density (target 2-3% primary) +8. **Validation:** All fields within character limits; primary keyword in title; no keyword stuffing (>5%); natural language preserved + +### Platform Character Limits + +| Field | Apple App Store | Google Play Store | +|-------|-----------------|-------------------| +| Title | 30 characters | 50 characters | +| Subtitle | 30 characters | N/A | +| Short Description | N/A | 80 characters | +| Keywords | 100 characters | N/A | +| Promotional Text | 170 characters | N/A | +| Full Description | 4,000 characters | 4,000 characters | +| What's New | 4,000 characters | 500 characters | + +### Description Structure + +``` +PARAGRAPH 1: Hook (50-100 words) +├── Address user pain point +├── State main value proposition +└── Include primary keyword + +PARAGRAPH 2-3: Features (100-150 words) +├── Top 5 features with benefits +├── Bullet points for scanability +└── Secondary keywords naturally integrated + +PARAGRAPH 4: Social Proof (50-75 words) +├── Download count or rating +├── Press mentions or awards +└── Summary of user testimonials + +PARAGRAPH 5: Call to Action (25-50 words) +├── Clear next step +└── Reassurance (free trial, no signup) +``` + +See: [references/platform-requirements.md](references/platform-requirements.md) + +--- + +## Competitor Analysis Workflow + +Analyze top competitors to identify keyword gaps and positioning opportunities. + +### Workflow: Analyze Competitor ASO Strategy + +1. Identify top 10 competitors: + - Direct competitors (same core function) + - Indirect competitors (overlapping audience) + - Category leaders (top downloads) +2. Extract competitor keywords from: + - App titles and subtitles + - First 100 words of descriptions + - Visible metadata patterns +3. Build competitor keyword matrix: + - Map which keywords each competitor targets + - Calculate coverage percentage per keyword +4. Identify keyword gaps: + - Keywords with <40% competitor coverage + - High volume terms competitors miss + - Long-tail opportunities +5. Analyze competitor visual assets: + - Icon design patterns + - Screenshot messaging and style + - Video presence and quality +6. Compare ratings and review patterns: + - Average rating by competitor + - Common praise themes + - Common complaint themes +7. Document positioning opportunities +8. **Validation:** 10+ competitors analyzed; keyword matrix complete; gaps identified with volume estimates; visual audit documented + +### Competitor Analysis Matrix + +| Analysis Area | Data Points | +|---------------|-------------| +| Keywords | Title keywords, description frequency | +| Metadata | Character utilization, keyword density | +| Visuals | Icon style, screenshot count/style | +| Ratings | Average rating, total count, velocity | +| Reviews | Top praise, top complaints | + +### Gap Analysis Template + +| Opportunity Type | Example | Action | +|------------------|---------|--------| +| Keyword gap | "habit tracker" (40% coverage) | Add to keyword field | +| Feature gap | Competitor lacks widget | Highlight in screenshots | +| Visual gap | No videos in top 5 | Create app preview | +| Messaging gap | None mention "free" | Test free positioning | + +--- + +## App Launch Workflow + +Execute a structured launch for maximum initial visibility. + +### Workflow: Launch App to Stores + +1. Complete pre-launch preparation (4 weeks before): + - Finalize keywords and metadata + - Prepare all visual assets + - Set up analytics (Firebase, Mixpanel) + - Build press kit and media list +2. Submit for review (2 weeks before): + - Complete all store requirements + - Verify compliance with guidelines + - Prepare launch communications +3. Configure post-launch systems: + - Set up review monitoring + - Prepare response templates + - Configure rating prompt timing +4. Execute launch day: + - Verify app is live in both stores + - Announce across all channels + - Begin review response cycle +5. Monitor initial performance (days 1-7): + - Track download velocity hourly + - Monitor reviews and respond within 24 hours + - Document any issues for quick fixes +6. Conduct 7-day retrospective: + - Compare performance to projections + - Identify quick optimization wins + - Plan first metadata update +7. Schedule first update (2 weeks post-launch) +8. **Validation:** App live in stores; analytics tracking; review responses within 24h; download velocity documented; first update scheduled + +### Pre-Launch Checklist + +| Category | Items | +|----------|-------| +| Metadata | Title, subtitle, description, keywords | +| Visual Assets | Icon, screenshots (all sizes), video | +| Compliance | Age rating, privacy policy, content rights | +| Technical | App binary, signing certificates | +| Analytics | SDK integration, event tracking | +| Marketing | Press kit, social content, email ready | + +### Launch Timing Considerations + +| Factor | Recommendation | +|--------|----------------| +| Day of week | Tuesday-Wednesday (avoid weekends) | +| Time of day | Morning in target market timezone | +| Seasonal | Align with relevant category seasons | +| Competition | Avoid major competitor launch dates | + +See: [references/aso-best-practices.md](references/aso-best-practices.md) + +--- + +## A/B Testing Workflow + +Test metadata and visual elements to improve conversion rates. + +### Workflow: Run A/B Test + +1. Select test element (prioritize by impact): + - Icon (highest impact) + - Screenshot 1 (high impact) + - Title (high impact) + - Short description (medium impact) +2. Form hypothesis: + ``` + If we [change], then [metric] will [improve/increase] by [amount] + because [rationale]. + ``` +3. Create variants: + - Control: Current version + - Treatment: Single variable change +4. Calculate required sample size: + - Baseline conversion rate + - Minimum detectable effect (usually 5%) + - Statistical significance (95%) +5. Launch test: + - Apple: Use Product Page Optimization + - Android: Use Store Listing Experiments +6. Run test for minimum duration: + - At least 7 days + - Until statistical significance reached +7. Analyze results: + - Compare conversion rates + - Check statistical significance + - Document learnings +8. **Validation:** Single variable tested; sample size sufficient; significance reached (95%); results documented; winner implemented + +### A/B Test Prioritization + +| Element | Conversion Impact | Test Complexity | +|---------|-------------------|-----------------| +| App Icon | 10-25% lift possible | Medium (design needed) | +| Screenshot 1 | 15-35% lift possible | Medium | +| Title | 5-15% lift possible | Low | +| Short Description | 5-10% lift possible | Low | +| Video | 10-20% lift possible | High | + +### Sample Size Quick Reference + +| Baseline CVR | Impressions Needed (per variant) | +|--------------|----------------------------------| +| 1% | 31,000 | +| 2% | 15,500 | +| 5% | 6,200 | +| 10% | 3,100 | + +### Test Documentation Template + +``` +TEST ID: ASO-2025-001 +ELEMENT: App Icon +HYPOTHESIS: A bolder color icon will increase conversion by 10% +START DATE: [Date] +END DATE: [Date] + +RESULTS: +├── Control CVR: 4.2% +├── Treatment CVR: 4.8% +├── Lift: +14.3% +├── Significance: 97% +└── Decision: Implement treatment + +LEARNINGS: +- Bold colors outperform muted tones in this category +- Apply to screenshot backgrounds for next test +``` + +--- + +## Before/After Examples + +### Title Optimization + +**Productivity App:** + +| Version | Title | Analysis | +|---------|-------|----------| +| Before | "MyTasks" | No keywords, brand only (8 chars) | +| After | "MyTasks - Todo List & Planner" | Primary + secondary keywords (29 chars) | + +**Fitness App:** + +| Version | Title | Analysis | +|---------|-------|----------| +| Before | "FitTrack Pro" | Generic modifier (12 chars) | +| After | "FitTrack: Workout Log & Gym" | Category keywords (27 chars) | + +### Subtitle Optimization (iOS) + +| Version | Subtitle | Analysis | +|---------|----------|----------| +| Before | "Get Things Done" | Vague, no keywords | +| After | "Daily Task Manager & Planner" | Two keywords, benefit clear | + +### Keyword Field Optimization (iOS) + +**Before (Inefficient - 89 chars, 8 keywords):** +``` +task manager, todo list, productivity app, daily planner, reminder app +``` + +**After (Optimized - 97 chars, 14 keywords):** +``` +task,todo,checklist,reminder,organize,daily,planner,schedule,deadline,goals,habit,widget,sync,team +``` + +**Improvements:** +- Removed spaces after commas (+8 chars) +- Removed duplicates (task manager → task) +- Removed plurals (reminders → reminder) +- Removed words in title +- Added more relevant keywords + +### Description Opening + +**Before:** +``` +MyTasks is a comprehensive task management solution designed +to help busy professionals organize their daily activities +and boost productivity. +``` + +**After:** +``` +Forget missed deadlines. MyTasks keeps every task, reminder, +and project in one place—so you focus on doing, not remembering. +Trusted by 500,000+ professionals. +``` + +**Improvements:** +- Leads with user pain point +- Specific benefit (not generic "boost productivity") +- Social proof included +- Keywords natural, not stuffed + +### Screenshot Caption Evolution + +| Version | Caption | Issue | +|---------|---------|-------| +| Before | "Task List Feature" | Feature-focused, passive | +| Better | "Create Task Lists" | Action verb, but still feature | +| Best | "Never Miss a Deadline" | Benefit-focused, emotional | + +--- + +## Tools and References + +### Scripts + +| Script | Purpose | Usage | +|--------|---------|-------| +| [keyword_analyzer.py](scripts/keyword_analyzer.py) | Analyze keywords for volume and competition | `python keyword_analyzer.py --keywords "todo,task,planner"` | +| [metadata_optimizer.py](scripts/metadata_optimizer.py) | Validate metadata character limits and density | `python metadata_optimizer.py --platform ios --title "App Title"` | +| [competitor_analyzer.py](scripts/competitor_analyzer.py) | Extract and compare competitor keywords | `python competitor_analyzer.py --competitors "App1,App2,App3"` | +| [aso_scorer.py](scripts/aso_scorer.py) | Calculate overall ASO health score | `python aso_scorer.py --app-id com.example.app` | +| [ab_test_planner.py](scripts/ab_test_planner.py) | Plan tests and calculate sample sizes | `python ab_test_planner.py --cvr 0.05 --lift 0.10` | +| [review_analyzer.py](scripts/review_analyzer.py) | Analyze review sentiment and themes | `python review_analyzer.py --app-id com.example.app` | +| [launch_checklist.py](scripts/launch_checklist.py) | Generate platform-specific launch checklists | `python launch_checklist.py --platform ios` | +| [localization_helper.py](scripts/localization_helper.py) | Manage multi-language metadata | `python localization_helper.py --locales "en,es,de,ja"` | + +### References + +| Document | Content | +|----------|---------| +| [platform-requirements.md](references/platform-requirements.md) | iOS and Android metadata specs, visual asset requirements | +| [aso-best-practices.md](references/aso-best-practices.md) | Optimization strategies, rating management, launch tactics | +| [keyword-research-guide.md](references/keyword-research-guide.md) | Research methodology, evaluation framework, tracking | + +### Assets + +| Template | Purpose | +|----------|---------| +| [aso-audit-template.md](assets/aso-audit-template.md) | Structured audit checklist for app store listings | + +--- + +## Platform Notes + +| Platform / Constraint | Behavior / Impact | +|-----------------------|-------------------| +| iOS keyword changes | Require app submission | +| iOS promotional text | Editable without an app update | +| Android metadata changes | Index in 1-2 hours | +| Android keyword field | None — use description instead | +| Keyword volume data | Estimates only; no official source | +| Competitor data | Public listings only | + +**When not to use this skill:** web apps (use web SEO), enterprise/internal apps, TestFlight-only betas, or paid advertising strategy. + +--- + +## Related Skills + +| Skill | Integration Point | +|-------|-------------------| +| [content-creator](../content-creator/) | App description copywriting | +| [marketing-demand-acquisition](../marketing-demand-acquisition/) | Launch promotion campaigns | +| [marketing-strategy-pmm](../marketing-strategy-pmm/) | Go-to-market planning | + +## Proactive Triggers + +- **No keyword optimization in title** → App title is the #1 ranking factor. Include top keyword. +- **Screenshots don't show value** → Screenshots should tell a story, not show UI. +- **No ratings strategy** → Below 4.0 stars kills conversion. Implement in-app rating prompts. +- **Description keyword-stuffed** → Natural language with keywords beats keyword stuffing. + +## Output Artifacts + +| When you ask for... | You get... | +|---------------------|------------| +| "ASO audit" | Full app store listing audit with prioritized fixes | +| "Keyword research" | Keyword list with search volume and difficulty scores | +| "Optimize my listing" | Rewritten title, subtitle, description, keyword field | + +## Communication + +All output passes quality verification: +- Self-verify: source attribution, assumption audit, confidence scoring +- Output format: Bottom Line → What (with confidence) → Why → How to Act +- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed. diff --git a/skills/app-store-optimization/_meta.json b/skills/app-store-optimization/_meta.json new file mode 100644 index 00000000..30bcf4ed --- /dev/null +++ b/skills/app-store-optimization/_meta.json @@ -0,0 +1,22 @@ +{ + "owner": "alirezarezvani", + "slug": "app-store-optimization", + "displayName": "App Store Optimization", + "latest": { + "version": "2.1.1", + "publishedAt": 1773070249169, + "commit": "https://github.com/openclaw/skills/commit/6bd4f910d29da812a68cc9a01f459f38ac63111d" + }, + "history": [ + { + "version": "1.0.0", + "publishedAt": 1770402531210, + "commit": "https://github.com/openclaw/skills/commit/792941c23326c617caa53dfd8c19be5e2ab88ae6" + }, + { + "version": "0.1.0", + "publishedAt": 1770028069136, + "commit": "https://github.com/clawdbot/skills/commit/d9a0182cc89c260cdec408409c1a65a87804d144" + } + ] +} diff --git a/skills/app-store-optimization/assets/aso-audit-template.md b/skills/app-store-optimization/assets/aso-audit-template.md new file mode 100644 index 00000000..4a727624 --- /dev/null +++ b/skills/app-store-optimization/assets/aso-audit-template.md @@ -0,0 +1,268 @@ +# ASO Audit Template + +Use this template to conduct a systematic App Store Optimization audit. + +--- + +## App Information + +| Field | Value | +|-------|-------| +| App Name | | +| Platform | [ ] iOS [ ] Android | +| Category | | +| Current Downloads | | +| Current Rating | | +| Audit Date | | + +--- + +## Metadata Audit + +### Title Analysis + +| Criterion | iOS (30 chars) | Android (50 chars) | +|-----------|----------------|---------------------| +| Current Title | | | +| Character Count | /30 | /50 | +| Primary Keyword Present | [ ] Yes [ ] No | [ ] Yes [ ] No | +| Brand Name Position | | | + +**Title Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +### Subtitle / Short Description + +| Criterion | iOS Subtitle (30 chars) | Android Short Desc (80 chars) | +|-----------|-------------------------|-------------------------------| +| Current Text | | | +| Character Count | /30 | /80 | +| Keywords Included | | | +| Benefit-Focused | [ ] Yes [ ] No | [ ] Yes [ ] No | + +**Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +### Keyword Field (iOS Only) + +| Criterion | Status | +|-----------|--------| +| Current Keywords | | +| Character Count | /100 | +| Duplicates Present | [ ] Yes [ ] No | +| Plurals Included | [ ] Yes [ ] No | +| Brand Names Included | [ ] Yes [ ] No | + +**Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +### Full Description + +| Criterion | iOS | Android | +|-----------|-----|---------| +| Character Count | /4000 | /4000 | +| Primary Keyword Density | % | % | +| Secondary Keywords (count) | | | +| Feature Bullets Present | [ ] Yes [ ] No | [ ] Yes [ ] No | +| Social Proof Included | [ ] Yes [ ] No | [ ] Yes [ ] No | +| CTA Present | [ ] Yes [ ] No | [ ] Yes [ ] No | + +**Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +--- + +## Visual Asset Audit + +### App Icon + +| Criterion | Status | +|-----------|--------| +| Recognizable at 60x60px | [ ] Yes [ ] No | +| Distinct from competitors | [ ] Yes [ ] No | +| Matches app design | [ ] Yes [ ] No | +| No text/words | [ ] Yes [ ] No | + +**Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +### Screenshots + +| Screenshot | Caption | Feature Shown | Score | +|------------|---------|---------------|-------| +| 1 (Hero) | | | /10 | +| 2 | | | /10 | +| 3 | | | /10 | +| 4 | | | /10 | +| 5 | | | /10 | + +| Criterion | Status | +|-----------|--------| +| Total Screenshots | /10 (iOS) or /8 (Android) | +| Captions Present | [ ] Yes [ ] No | +| Consistent Style | [ ] Yes [ ] No | +| First 3 Show Value | [ ] Yes [ ] No | +| Device Frames Used | [ ] Yes [ ] No | + +**Overall Screenshot Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +### App Preview Video + +| Criterion | Status | +|-----------|--------| +| Video Present | [ ] Yes [ ] No | +| Duration | seconds | +| Shows Core Features | [ ] Yes [ ] No | +| Hook in First 5 Seconds | [ ] Yes [ ] No | +| CTA at End | [ ] Yes [ ] No | + +**Score:** ___/10 + +--- + +## Keyword Performance Audit + +### Current Keyword Rankings + +| Keyword | Current Rank | Volume | Competition | Score | +|---------|--------------|--------|-------------|-------| +| | | | | | +| | | | | | +| | | | | | +| | | | | | +| | | | | | + +### Keyword Opportunities + +| Keyword | Current Rank | Potential | Action | +|---------|--------------|-----------|--------| +| | | | | +| | | | | +| | | | | + +--- + +## Rating & Review Audit + +### Rating Summary + +| Metric | Value | +|--------|-------| +| Current Average Rating | /5.0 | +| Total Ratings | | +| Ratings (Last 30 Days) | | +| 5-Star Percentage | % | +| 1-Star Percentage | % | + +### Review Analysis + +| Category | Count | Common Themes | +|----------|-------|---------------| +| Positive (4-5 stars) | | | +| Neutral (3 stars) | | | +| Negative (1-2 stars) | | | + +### Response Rate + +| Metric | Value | +|--------|-------| +| Reviews Responded | % | +| Avg Response Time | hours | + +**Rating Score:** ___/10 + +**Recommendations:** +- [ ] +- [ ] + +--- + +## Competitor Comparison + +### Top 3 Competitors + +| Metric | Your App | Competitor 1 | Competitor 2 | Competitor 3 | +|--------|----------|--------------|--------------|--------------| +| Name | | | | | +| Rating | | | | | +| Total Ratings | | | | | +| Downloads | | | | | +| Title Keywords | | | | | +| Screenshot Count | | | | | + +### Competitive Gaps + +| Gap Identified | Opportunity | +|----------------|-------------| +| | | +| | | +| | | + +--- + +## Overall ASO Score + +| Category | Weight | Score | Weighted | +|----------|--------|-------|----------| +| Title/Metadata | 25% | /10 | | +| Keywords | 25% | /10 | | +| Visual Assets | 25% | /10 | | +| Ratings/Reviews | 25% | /10 | | +| **TOTAL** | 100% | | **/100** | + +--- + +## Priority Action Items + +### High Priority (This Week) + +1. [ ] +2. [ ] +3. [ ] + +### Medium Priority (This Month) + +1. [ ] +2. [ ] +3. [ ] + +### Low Priority (This Quarter) + +1. [ ] +2. [ ] +3. [ ] + +--- + +## Audit Sign-Off + +| Role | Name | Date | +|------|------|------| +| Auditor | | | +| Reviewer | | | +| App Owner | | | + +--- + +## Notes + +_Additional observations and context:_ diff --git a/skills/app-store-optimization/expected_output.json b/skills/app-store-optimization/expected_output.json new file mode 100644 index 00000000..9832693a --- /dev/null +++ b/skills/app-store-optimization/expected_output.json @@ -0,0 +1,170 @@ +{ + "request_type": "keyword_research", + "app_name": "TaskFlow Pro", + "keyword_analysis": { + "total_keywords_analyzed": 25, + "primary_keywords": [ + { + "keyword": "task manager", + "search_volume": 45000, + "competition_level": "high", + "relevance_score": 0.95, + "difficulty_score": 72.5, + "potential_score": 78.3, + "recommendation": "High priority - target immediately" + }, + { + "keyword": "productivity app", + "search_volume": 38000, + "competition_level": "high", + "relevance_score": 0.90, + "difficulty_score": 68.2, + "potential_score": 75.1, + "recommendation": "High priority - target immediately" + }, + { + "keyword": "todo list", + "search_volume": 52000, + "competition_level": "very_high", + "relevance_score": 0.85, + "difficulty_score": 78.9, + "potential_score": 71.4, + "recommendation": "High priority - target immediately" + } + ], + "secondary_keywords": [ + { + "keyword": "team task manager", + "search_volume": 8500, + "competition_level": "medium", + "relevance_score": 0.88, + "difficulty_score": 42.3, + "potential_score": 68.7, + "recommendation": "Good opportunity - include in metadata" + }, + { + "keyword": "project planning app", + "search_volume": 12000, + "competition_level": "medium", + "relevance_score": 0.75, + "difficulty_score": 48.1, + "potential_score": 64.2, + "recommendation": "Good opportunity - include in metadata" + } + ], + "long_tail_keywords": [ + { + "keyword": "ai task prioritization", + "search_volume": 2800, + "competition_level": "low", + "relevance_score": 0.95, + "difficulty_score": 25.4, + "potential_score": 82.6, + "recommendation": "Excellent long-tail opportunity" + }, + { + "keyword": "team productivity tool", + "search_volume": 3500, + "competition_level": "low", + "relevance_score": 0.85, + "difficulty_score": 28.7, + "potential_score": 79.3, + "recommendation": "Excellent long-tail opportunity" + } + ] + }, + "competitor_insights": { + "competitors_analyzed": 4, + "common_keywords": [ + "task", + "todo", + "list", + "productivity", + "organize", + "manage" + ], + "keyword_gaps": [ + { + "keyword": "ai prioritization", + "used_by": ["None of the major competitors"], + "opportunity": "Unique positioning opportunity" + }, + { + "keyword": "smart task manager", + "used_by": ["Things 3"], + "opportunity": "Underutilized by most competitors" + } + ] + }, + "metadata_recommendations": { + "apple_app_store": { + "title_options": [ + { + "title": "TaskFlow - AI Task Manager", + "length": 26, + "keywords_included": ["task manager", "ai"], + "strategy": "brand_plus_primary" + }, + { + "title": "TaskFlow: Smart Todo & Tasks", + "length": 29, + "keywords_included": ["todo", "tasks"], + "strategy": "brand_plus_multiple" + } + ], + "subtitle_recommendation": "AI-Powered Team Productivity", + "keyword_field": "productivity,organize,planner,schedule,workflow,reminders,collaboration,calendar,sync,priorities", + "description_focus": "Lead with AI differentiation, emphasize team features" + }, + "google_play_store": { + "title_options": [ + { + "title": "TaskFlow - AI Task Manager & Team Productivity", + "length": 48, + "keywords_included": ["task manager", "ai", "team", "productivity"], + "strategy": "keyword_rich" + } + ], + "short_description_recommendation": "AI task manager - Organize, prioritize, and collaborate with your team", + "description_focus": "Keywords naturally integrated throughout 4000 character description" + } + }, + "strategic_recommendations": [ + "Focus on 'AI prioritization' as unique differentiator - low competition, high relevance", + "Target 'team task manager' and 'team productivity' keywords - good search volume, lower competition than generic terms", + "Include long-tail keywords in description for additional discovery opportunities", + "Test title variations with A/B testing after launch", + "Monitor competitor keyword changes quarterly" + ], + "priority_actions": [ + { + "action": "Optimize app title with primary keyword", + "priority": "high", + "expected_impact": "15-25% improvement in search visibility" + }, + { + "action": "Create description highlighting AI features with natural keyword integration", + "priority": "high", + "expected_impact": "10-15% improvement in conversion rate" + }, + { + "action": "Plan A/B tests for icon and screenshots post-launch", + "priority": "medium", + "expected_impact": "5-10% improvement in conversion rate" + } + ], + "aso_health_estimate": { + "current_score": "N/A (pre-launch)", + "potential_score_with_optimizations": "75-80/100", + "key_strengths": [ + "Unique AI differentiation", + "Clear target audience", + "Strong feature set" + ], + "areas_to_develop": [ + "Build rating volume post-launch", + "Monitor and respond to reviews", + "Continuous keyword optimization" + ] + } +} diff --git a/skills/app-store-optimization/references/aso-best-practices.md b/skills/app-store-optimization/references/aso-best-practices.md new file mode 100644 index 00000000..9ce7ca24 --- /dev/null +++ b/skills/app-store-optimization/references/aso-best-practices.md @@ -0,0 +1,403 @@ +# ASO Best Practices Reference + +Optimization strategies for improving app store visibility, conversion, and rankings. + +--- + +## Table of Contents + +- [Keyword Optimization](#keyword-optimization) +- [Metadata Optimization](#metadata-optimization) +- [Visual Asset Optimization](#visual-asset-optimization) +- [Rating and Review Management](#rating-and-review-management) +- [Launch Strategy](#launch-strategy) +- [A/B Testing Framework](#ab-testing-framework) +- [Conversion Optimization](#conversion-optimization) +- [Common Mistakes to Avoid](#common-mistakes-to-avoid) + +--- + +## Keyword Optimization + +### Keyword Research Process + +1. **Brainstorm seed keywords** - Core terms users search for +2. **Expand with variations** - Synonyms, related terms, long-tail +3. **Analyze competition** - Check difficulty scores +4. **Evaluate search volume** - Prioritize high-volume terms +5. **Test and iterate** - Monitor rankings and adjust + +### Keyword Selection Criteria + +| Factor | Weight | Evaluation Method | +|--------|--------|-------------------| +| Relevance | 40% | Does it describe app function? | +| Search Volume | 30% | Monthly search estimates | +| Competition | 20% | Number of ranking apps | +| Conversion Potential | 10% | User intent alignment | + +### Keyword Placement Priority + +| Location | Search Weight | Example | +|----------|---------------|---------| +| App Title | Highest | "TaskMaster - Todo List Manager" | +| Subtitle (iOS) | High | "Organize Your Daily Tasks" | +| Keyword Field (iOS) | High | "planner,reminder,checklist" | +| Short Description (Android) | High | "Simple task manager for busy professionals" | +| Full Description | Medium | Natural keyword usage throughout | + +### Long-Tail Keyword Strategy + +Long-tail keywords have lower volume but higher conversion: + +| Type | Example | Volume | Competition | Conversion | +|------|---------|--------|-------------|------------| +| Short-tail | "todo app" | High | High | Low | +| Mid-tail | "daily task manager" | Medium | Medium | Medium | +| Long-tail | "free todo list with reminders" | Low | Low | High | + +**Formula for keyword priority:** +``` +Score = (Volume × 0.3) + (1/Competition × 0.3) + (Relevance × 0.4) +``` + +--- + +## Metadata Optimization + +### Title Optimization + +**Structure Formula:** +``` +[Brand Name] - [Primary Keyword] [Secondary Keyword/Benefit] +``` + +**Examples by category:** + +| Category | Before | After | +|----------|--------|-------| +| Productivity | "MyTasks" | "MyTasks - Todo List & Planner" | +| Fitness | "FitTrack" | "FitTrack: Workout & Gym Log" | +| Finance | "MoneyApp" | "MoneyApp - Budget Tracker" | +| Photo | "SnapEdit" | "SnapEdit: Photo Editor & AI" | + +**Title Optimization Checklist:** +- [ ] Primary keyword within first 3 words +- [ ] Brand name is memorable and unique +- [ ] Character count matches platform limit +- [ ] No keyword stuffing +- [ ] Readable and natural sounding + +### Description Optimization + +**Full Description Structure:** + +``` +PARAGRAPH 1: Hook + Primary Benefit (50-100 words) +- Address user pain point +- State main value proposition +- Include primary keyword naturally + +PARAGRAPH 2-3: Feature Highlights (100-150 words) +- Top 3-5 features with benefits +- Use bullet points or emojis for scanability +- Include secondary keywords + +PARAGRAPH 4: Social Proof (50-75 words) +- Download numbers or ratings +- Press mentions or awards +- User testimonials (summarized) + +PARAGRAPH 5: Call to Action (25-50 words) +- Clear next step +- Urgency or incentive +- Reassurance (free trial, no credit card) +``` + +**Keyword Density Target:** +- Primary keyword: 2-3% (8-12 mentions in 4000 chars) +- Secondary keywords: 1-2% each (4-8 mentions each) + +### Subtitle Optimization (iOS) + +**Effective Subtitle Formulas:** + +| Formula | Example | +|---------|---------| +| [Verb] + [Benefit] | "Organize Your Life" | +| [Adjective] + [Category] | "Smart Task Manager" | +| [Feature] + [Feature] | "Lists, Reminders & Notes" | +| [Audience] + [Solution] | "For Busy Professionals" | + +--- + +## Visual Asset Optimization + +### App Icon Best Practices + +| Principle | Do | Don't | +|-----------|-----|-------| +| Simplicity | Single focal element | Multiple competing elements | +| Recognizability | Works at 60x60px | Requires large size to read | +| Uniqueness | Distinct from competitors | Generic category icon | +| Color | Bold, contrasting colors | Muted or similar to background | +| Text | None or single letter | Full words or app name | + +**Icon Testing Questions:** +1. Is it recognizable at 29x29px (smallest iOS size)? +2. Does it stand out in search results? +3. Does it communicate app function? +4. Is it distinct from top 10 category competitors? + +### Screenshot Optimization + +**Screenshot Hierarchy:** + +| Position | Purpose | Content Strategy | +|----------|---------|------------------| +| Screenshot 1 | Hook/Hero | Main value proposition + key UI | +| Screenshot 2 | Primary Feature | Most-used feature demonstration | +| Screenshot 3 | Secondary Feature | Differentiating capability | +| Screenshot 4 | Social Proof | Ratings, awards, user count | +| Screenshot 5+ | Additional Features | Supporting functionality | + +**Caption Best Practices:** +- Maximum 5-7 words per caption +- Action-oriented verbs ("Track", "Organize", "Discover") +- Benefit-focused, not feature-focused +- Consistent typography and style + +**Example Caption Evolution:** + +| Weak | Better | Best | +|------|--------|------| +| "Task List Feature" | "Create Task Lists" | "Never Forget a Task Again" | +| "Calendar View" | "See Your Schedule" | "Plan Your Week in Seconds" | +| "Notifications" | "Get Reminders" | "Stay on Top of Deadlines" | + +### Video Preview Strategy + +**Video Structure (30 seconds):** + +| Seconds | Content | +|---------|---------| +| 0-5 | Hook: Show end result or main benefit | +| 5-15 | Demo: Core feature in action | +| 15-25 | Features: Quick feature montage | +| 25-30 | CTA: Logo and download prompt | + +--- + +## Rating and Review Management + +### Review Response Framework + +**For Negative Reviews (1-2 stars):** + +``` +Structure: +1. Acknowledge the issue (1 sentence) +2. Apologize without making excuses (1 sentence) +3. Offer solution or next step (1-2 sentences) +4. Invite direct contact (1 sentence) + +Example: +"We're sorry the syncing issues are affecting your experience. +Our team is actively working on a fix for the next update. +In the meantime, please try logging out and back in, which +resolves this for most users. If issues persist, email us at +support@app.com and we'll prioritize your case." +``` + +**For Positive Reviews (4-5 stars):** + +``` +Structure: +1. Thank sincerely (1 sentence) +2. Acknowledge specific praise (1 sentence) +3. Encourage continued use or sharing (1 sentence) + +Example: +"Thank you for the kind words! We're thrilled the reminder +feature helps you stay organized. If you're enjoying the app, +we'd love if you'd share it with friends who might benefit." +``` + +### Rating Improvement Tactics + +| Tactic | Implementation | Expected Impact | +|--------|----------------|-----------------| +| In-app prompt timing | After positive action (task completed, milestone reached) | +0.3 stars | +| Bug fix velocity | Address 1-star issues within 7 days | +0.2 stars | +| Response rate | Reply to 80%+ of reviews | +0.1 stars | +| Feature requests | Implement top-requested features | +0.2 stars | + +### Review Prompt Best Practices + +**When to prompt:** +- After user completes 5+ successful sessions +- After milestone achievement (first task completed, 7-day streak) +- After positive in-app feedback ("Was this helpful? Yes") + +**When NOT to prompt:** +- First session +- After error or crash +- During critical workflow +- More than once per 30 days + +--- + +## Launch Strategy + +### Pre-Launch Checklist + +**4 Weeks Before Launch:** +- [ ] Finalize app name and keywords +- [ ] Complete all metadata fields +- [ ] Prepare all visual assets +- [ ] Set up analytics (Firebase, Mixpanel) +- [ ] Create press kit and media assets +- [ ] Build email list for launch notification + +**2 Weeks Before Launch:** +- [ ] Submit for app review +- [ ] Prepare social media content +- [ ] Brief press and influencers +- [ ] Set up review response templates +- [ ] Configure in-app rating prompts + +**Launch Day:** +- [ ] Verify app is live in stores +- [ ] Announce across all channels +- [ ] Monitor reviews and respond quickly +- [ ] Track download velocity +- [ ] Document any issues for Day 2 fix + +### Update Cadence + +| Update Type | Frequency | ASO Impact | +|-------------|-----------|------------| +| Bug fixes | As needed | Prevents rating drops | +| Minor features | Every 2-4 weeks | Maintains freshness signal | +| Major features | Every 4-8 weeks | Opportunity for "What's New" | +| Metadata refresh | Every 4-6 weeks | Keyword optimization cycle | + +### Seasonal Optimization + +| Season | Optimization Focus | Example Categories | +|--------|--------------------|--------------------| +| Jan (New Year) | Resolutions, goals | Fitness, Productivity | +| Feb (Valentine's) | Dating, relationships | Dating, Photo | +| Mar-Apr (Tax) | Finance, organization | Finance, Productivity | +| May-Jun (Summer) | Travel, fitness | Travel, Health | +| Aug-Sep (Back to School) | Education, organization | Education, Productivity | +| Nov-Dec (Holidays) | Shopping, social | Shopping, Social | + +--- + +## A/B Testing Framework + +### Test Prioritization Matrix + +| Element | Impact | Ease | Priority | +|---------|--------|------|----------| +| App Icon | High | Medium | 1 | +| Screenshot 1 | High | Medium | 2 | +| Title | High | Easy | 3 | +| Short Description | Medium | Easy | 4 | +| Screenshots 2-5 | Medium | Medium | 5 | +| Video | Medium | Hard | 6 | + +### Sample Size Calculator + +**Formula:** +``` +Sample Size = (2 × (Z² × p × (1-p))) / E² + +Where: +Z = 1.96 (for 95% confidence) +p = baseline conversion rate +E = minimum detectable effect (usually 0.05) +``` + +**Quick Reference:** + +| Baseline CVR | Min. Impressions for 5% Lift | +|--------------|------------------------------| +| 1% | 31,000 per variant | +| 2% | 15,500 per variant | +| 5% | 6,200 per variant | +| 10% | 3,100 per variant | + +### Test Duration Guidelines + +| Daily Impressions | Minimum Test Duration | +|-------------------|----------------------| +| 1,000+ | 7 days | +| 500-1,000 | 14 days | +| 100-500 | 30 days | +| <100 | Not recommended | + +--- + +## Conversion Optimization + +### Conversion Funnel Metrics + +| Stage | Metric | Benchmark | +|-------|--------|-----------| +| Discovery | Impressions | Category dependent | +| Consideration | Page Views | 30-50% of impressions | +| Conversion | Installs | 3-8% of page views | +| Activation | First Open | 70-90% of installs | + +### Conversion Optimization Levers + +| Lever | Typical Lift | Effort | +|-------|--------------|--------| +| Icon redesign | 10-25% | High | +| Screenshot optimization | 15-35% | Medium | +| Title keyword optimization | 5-15% | Low | +| Description rewrite | 5-10% | Low | +| Video addition | 10-20% | High | +| Localization | 20-50% per market | Medium | + +--- + +## Common Mistakes to Avoid + +### Keyword Mistakes + +| Mistake | Problem | Solution | +|---------|---------|----------| +| Keyword stuffing | Spam detection, rejection | Natural usage, 2-3% density | +| Competitor names | Guideline violation | Focus on category terms | +| Duplicate keywords | Wasted character space | Remove duplicates from keyword field | +| Ignoring long-tail | Missing conversion | Include specific phrases | + +### Metadata Mistakes + +| Mistake | Problem | Solution | +|---------|---------|----------| +| Vague descriptions | Low conversion | Specific benefits and features | +| Feature-focused copy | Doesn't resonate | Benefit-focused messaging | +| Outdated information | Misleading users | Update with each release | +| Missing localization | Lost global revenue | Prioritize top 5 markets | + +### Visual Asset Mistakes + +| Mistake | Problem | Solution | +|---------|---------|----------| +| Text-heavy screenshots | Unreadable on phones | Minimal text, clear UI focus | +| Inconsistent style | Unprofessional appearance | Design system for all assets | +| Portrait-only screenshots | Missed tablet users | Include landscape variants | +| No social proof | Lower trust | Add ratings, awards, press | + +### Launch Mistakes + +| Mistake | Problem | Solution | +|---------|---------|----------| +| Launching on Friday | No support over weekend | Launch Tuesday-Wednesday | +| No analytics setup | Can't measure success | Firebase/Mixpanel before launch | +| Immediate rating prompt | Negative ratings | Wait for positive experience | +| Ignoring reviews | Declining ratings | Respond within 24-48 hours | diff --git a/skills/app-store-optimization/references/keyword-research-guide.md b/skills/app-store-optimization/references/keyword-research-guide.md new file mode 100644 index 00000000..1110aa65 --- /dev/null +++ b/skills/app-store-optimization/references/keyword-research-guide.md @@ -0,0 +1,419 @@ +# Keyword Research Guide + +Systematic approach to discovering, evaluating, and selecting keywords for app store optimization. + +--- + +## Table of Contents + +- [Keyword Research Methodology](#keyword-research-methodology) +- [Keyword Evaluation Framework](#keyword-evaluation-framework) +- [Competitor Keyword Analysis](#competitor-keyword-analysis) +- [Keyword Mapping Strategy](#keyword-mapping-strategy) +- [Keyword Tracking and Iteration](#keyword-tracking-and-iteration) + +--- + +## Keyword Research Methodology + +### Phase 1: Seed Keyword Generation + +Start by generating initial keyword ideas from multiple sources. + +**Source 1: Core App Functions** + +List every action or problem the app solves: + +``` +Example for a task management app: +- Create tasks +- Set reminders +- Track deadlines +- Organize projects +- Collaborate with team +- Plan daily schedule +``` + +**Source 2: User Language Mapping** + +Match developer terminology to user searches: + +| Developer Term | User Search Terms | +|----------------|-------------------| +| Task management | todo list, task app, tasks | +| Project organization | project planner, project tracker | +| Deadline tracking | due date reminder, deadline app | +| Time blocking | schedule planner, calendar app | +| GTD methodology | getting things done, productivity system | + +**Source 3: App Store Autocomplete** + +Type seed keywords into App Store/Play Store search and record suggestions: + +``` +"todo" → todo list, todo app, todo list app, todolist widget +"task" → task manager, task planner, task list, tasks to do +"remind" → reminder app, reminder, reminders widget, remind me +``` + +**Source 4: Competitor Analysis** + +Extract keywords from top 10 competitors in category (detailed in section below). + +### Phase 2: Keyword Expansion + +**Expansion Techniques:** + +| Technique | Example (seed: "todo") | +|-----------|------------------------| +| Add modifiers | free todo, best todo, simple todo | +| Add actions | make todo list, create todo, organize todo | +| Add platforms | todo app iphone, todo for mac, todo widget | +| Add audiences | todo for students, business todo, family todo | +| Add features | todo with reminders, todo calendar, todo sync | +| Add problems | forgot tasks todo, procrastination todo | + +**Keyword Matrix Template:** + +| Core Term | Modifier 1 | Modifier 2 | Full Keyword | +|-----------|------------|------------|--------------| +| todo | free | app | free todo app | +| todo | best | iphone | best todo iphone | +| task | manager | simple | simple task manager | +| reminder | daily | widget | daily reminder widget | +| planner | weekly | calendar | weekly planner calendar | + +### Phase 3: Keyword Filtering + +Remove irrelevant or low-quality keywords: + +**Exclusion Criteria:** + +| Criterion | Reason | Example | +|-----------|--------|---------| +| Competitor brand names | Policy violation | "todoist alternative" | +| Unrelated categories | Low conversion | "todo games" | +| Plural duplicates (iOS) | Wasted space | "tasks" when "task" exists | +| Single characters | No search value | "to do" vs "todo" | + +--- + +## Keyword Evaluation Framework + +### Keyword Scoring Model + +Evaluate each keyword on four dimensions: + +**1. Search Volume (0-100)** + +| Volume Level | Score | Monthly Searches | +|--------------|-------|------------------| +| Very High | 80-100 | 50,000+ | +| High | 60-79 | 10,000-49,999 | +| Medium | 40-59 | 1,000-9,999 | +| Low | 20-39 | 100-999 | +| Very Low | 0-19 | <100 | + +**2. Competition (0-100, inverted)** + +| Competition | Score | Top 10 App Ratings | +|-------------|-------|-------------------| +| Very Low | 80-100 | Average <4.0 stars | +| Low | 60-79 | Average 4.0-4.2 stars | +| Medium | 40-59 | Average 4.3-4.5 stars | +| High | 20-39 | Average 4.6-4.8 stars | +| Very High | 0-19 | Average 4.9+ stars | + +**3. Relevance (0-100)** + +| Relevance | Score | Criteria | +|-----------|-------|----------| +| Exact Match | 90-100 | Keyword describes core function | +| Strong Match | 70-89 | Keyword describes major feature | +| Moderate Match | 50-69 | Keyword describes secondary feature | +| Weak Match | 30-49 | Keyword tangentially related | +| No Match | 0-29 | Keyword unrelated to app | + +**4. Conversion Potential (0-100)** + +| Intent | Score | User Query Type | +|--------|-------|-----------------| +| Transactional | 80-100 | "best [app type]", "[app type] app" | +| Commercial | 60-79 | "free [app type]", "[app type] for [use]" | +| Informational | 40-59 | "how to [action]", "what is [concept]" | +| Navigational | 20-39 | "[brand name]", "[specific app]" | + +### Composite Score Calculation + +``` +Keyword Score = (Volume × 0.25) + (Competition × 0.25) + + (Relevance × 0.35) + (Conversion × 0.15) +``` + +**Score Interpretation:** + +| Score Range | Priority | Action | +|-------------|----------|--------| +| 80-100 | Primary | Target in title and keyword field | +| 60-79 | Secondary | Include in subtitle/description | +| 40-59 | Tertiary | Use in long description only | +| 0-39 | Deprioritize | Do not target | + +### Keyword Evaluation Worksheet + +``` +KEYWORD EVALUATION + +Keyword: "task manager app" +Date: [Date] + +SCORES: +├── Search Volume: 72/100 (High - ~25,000/month) +├── Competition: 45/100 (Medium - 4.4 avg rating in top 10) +├── Relevance: 95/100 (Exact match to core function) +└── Conversion: 85/100 (Transactional intent) + +COMPOSITE SCORE: 74.5/100 + +RECOMMENDATION: Secondary Priority +- Include in subtitle or short description +- Not competitive enough for title (dominated by Todoist, Any.do) +- Consider long-tail variant: "simple task manager app" +``` + +--- + +## Competitor Keyword Analysis + +### Competitor Identification + +**Step 1: Direct Competitors** +Apps solving the same problem for the same audience. + +**Step 2: Indirect Competitors** +Apps solving related problems or targeting overlapping audiences. + +**Step 3: Category Leaders** +Top 10-20 apps by downloads in primary category. + +### Competitor Keyword Extraction + +**From App Title:** +``` +Competitor: "Todoist: To-Do List & Tasks" +Keywords: todoist, to-do list, tasks, to do +``` + +**From Subtitle (iOS):** +``` +Competitor subtitle: "Task Manager & Planner" +Keywords: task manager, planner +``` + +**From Description (First 100 words):** +Identify frequently used terms: +``` +"Todoist is the world's favorite task manager and to-do list app. +Organize work and life, hit your goals, and find productivity..." + +Extracted: task manager, to-do list, organize, goals, productivity +``` + +### Competitor Keyword Matrix + +| Keyword | Comp 1 | Comp 2 | Comp 3 | Comp 4 | Comp 5 | Coverage | +|---------|--------|--------|--------|--------|--------|----------| +| task manager | ✓ | ✓ | ✓ | ✓ | ✓ | 100% | +| to-do list | ✓ | ✓ | ✓ | ✓ | | 80% | +| planner | ✓ | ✓ | | ✓ | ✓ | 80% | +| reminder | ✓ | ✓ | ✓ | | | 60% | +| productivity | ✓ | | ✓ | ✓ | | 60% | +| checklist | | ✓ | | ✓ | ✓ | 60% | +| project | ✓ | ✓ | | | | 40% | +| habit | | | ✓ | | ✓ | 40% | + +**Analysis:** +- 100% coverage = Highly competitive, essential keyword +- 60-80% coverage = Important category term +- 40% coverage = Potential differentiator +- <40% coverage = Unique opportunity or irrelevant + +### Keyword Gap Analysis + +Identify keywords competitors miss: + +``` +KEYWORD GAP ANALYSIS + +Underserved Keywords (Low competitor coverage, decent volume): +1. "daily planner widget" - 2/10 competitors, 5,000 searches +2. "task list for teams" - 3/10 competitors, 3,500 searches +3. "todo with calendar sync" - 1/10 competitors, 2,800 searches + +Opportunity Assessment: +- "daily planner widget" → Add widget feature, target keyword +- "task list for teams" → Already have feature, update metadata +- "todo with calendar sync" → Feature gap, add to roadmap +``` + +--- + +## Keyword Mapping Strategy + +### Keyword Placement Map + +Assign each keyword to specific metadata locations: + +``` +KEYWORD PLACEMENT MAP + +PRIMARY (Title + Keyword Field): +├── task manager (Score: 82) +├── todo list (Score: 78) +└── planner (Score: 75) + +SECONDARY (Subtitle + Short Description): +├── reminder app (Score: 68) +├── daily tasks (Score: 65) +└── organize (Score: 62) + +TERTIARY (Full Description): +├── checklist (Score: 55) +├── productivity (Score: 52) +├── schedule (Score: 48) +├── deadline (Score: 45) +└── project management (Score: 42) +``` + +### iOS Keyword Field Strategy + +**100 Character Optimization:** + +``` +STEP 1: List all target keywords +task,manager,todo,list,planner,reminder,organize,daily,checklist, +productivity,schedule,deadline,project,goals,habit,widget,sync, +team,collaborate,notes,calendar + +STEP 2: Remove duplicates from title +Title: "TaskFlow - Todo List Manager" +Remove: task, todo, list, manager + +STEP 3: Remove plurals +Keep: reminder (not reminders) +Keep: goal (not goals) + +STEP 4: Prioritize by score and fit +Final 100 chars: +planner,reminder,organize,daily,checklist,productivity,schedule, +deadline,project,goals,habit,widget,sync,team,collaborate + +Character count: 98/100 +``` + +### Android Description Keyword Integration + +**Natural keyword placement in 4,000 characters:** + +``` +PARAGRAPH 1 (Hook - 300 chars): +Keywords: task manager, todo list, organize +"TaskFlow is the task manager trusted by 2 million users. Create +your perfect todo list and organize everything that matters..." + +PARAGRAPH 2 (Features - 800 chars): +Keywords: reminder, checklist, deadline, project +"Set smart reminders that notify you at the right time. Build +checklists for any project. Never miss a deadline with..." + +PARAGRAPH 3 (Benefits - 600 chars): +Keywords: productivity, schedule, goals +"Boost your productivity with proven planning methods. Schedule +your day in minutes. Track goals and celebrate..." + +PARAGRAPH 4 (Differentiators - 500 chars): +Keywords: widget, sync, team, collaborate +"Beautiful widgets keep tasks visible. Sync across all devices +instantly. Invite your team to collaborate on..." + +Total keyword coverage: 14 keywords naturally integrated +``` + +--- + +## Keyword Tracking and Iteration + +### Ranking Tracking Cadence + +| Frequency | Action | +|-----------|--------| +| Daily | Track top 5-10 primary keywords | +| Weekly | Full keyword set review | +| Monthly | Competitor keyword comparison | +| Quarterly | Full keyword research refresh | + +### Keyword Performance Metrics + +| Metric | Target | Action if Below | +|--------|--------|-----------------| +| Top 10 ranking | 3+ keywords | Increase keyword weight | +| Top 50 ranking | 10+ keywords | Maintain current strategy | +| Ranking velocity | Improving trend | Continue optimization | +| Conversion rate | >5% | Review relevance alignment | + +### Iteration Process + +**Monthly Keyword Audit:** + +``` +1. EXPORT current rankings + - List all tracked keywords + - Record current position + - Note 30-day trend (up/down/stable) + +2. IDENTIFY opportunities + - Keywords improving but not top 10 + - Keywords declining from previous position + - New high-volume keywords in category + +3. PRIORITIZE changes + - Boost: Keywords at position 11-20 + - Maintain: Keywords at position 1-10 + - Replace: Keywords at position 50+ with no improvement + +4. IMPLEMENT updates + - Adjust keyword field (iOS) + - Update description (Android) + - Modify subtitle if needed + +5. DOCUMENT changes + - Record what changed and why + - Set reminder for 2-week check-in +``` + +### Keyword Testing Log Template + +``` +KEYWORD TEST LOG + +Test ID: KW-2025-001 +Date Started: [Date] +Keywords Changed: + - Added: "habit tracker" (replacing "goals app") + - Added: "daily routine" (replacing "schedule planner") + +Rationale: +- "habit tracker" has 3x volume of "goals app" +- "daily routine" trending up 40% in category + +Baseline Rankings: +- "habit tracker": Not ranked +- "daily routine": Position 87 + +30-Day Results: +- "habit tracker": Position 34 (+53) +- "daily routine": Position 28 (+59) + +Conclusion: Test successful - retain new keywords +Next Action: Target subtitle position for "habit tracker" +``` diff --git a/skills/app-store-optimization/references/platform-requirements.md b/skills/app-store-optimization/references/platform-requirements.md new file mode 100644 index 00000000..16a08045 --- /dev/null +++ b/skills/app-store-optimization/references/platform-requirements.md @@ -0,0 +1,324 @@ +# Platform Requirements Reference + +Technical specifications and metadata requirements for Apple App Store and Google Play Store. + +--- + +## Table of Contents + +- [Apple App Store Requirements](#apple-app-store-requirements) +- [Google Play Store Requirements](#google-play-store-requirements) +- [Visual Asset Specifications](#visual-asset-specifications) +- [Localization Requirements](#localization-requirements) +- [Compliance Guidelines](#compliance-guidelines) + +--- + +## Apple App Store Requirements + +### Metadata Character Limits + +| Field | Character Limit | Notes | +|-------|----------------|-------| +| App Name (Title) | 30 characters | Visible in search results | +| Subtitle | 30 characters | iOS 11+ only, appears below title | +| Promotional Text | 170 characters | Editable without app update | +| Description | 4,000 characters | Not indexed for search | +| Keywords Field | 100 characters | Comma-separated, no spaces after commas | +| What's New | 4,000 characters | Release notes for updates | +| Developer Name | 255 characters | Company or individual name | +| Support URL | Required | Must be valid HTTPS URL | +| Privacy Policy URL | Required | Must be valid HTTPS URL | + +### Keyword Field Optimization Rules + +1. **No duplicates** - Words in title are already indexed +2. **No plurals** - Apple indexes both singular and plural forms +3. **No spaces after commas** - Wastes character space +4. **No brand names** - Violates App Store guidelines +5. **No category names** - Already indexed via category selection + +**Example - Efficient keyword field:** +``` +task,todo,checklist,reminder,productivity,organize,schedule,planner,goals,habit +``` + +**Example - Inefficient keyword field (avoid):** +``` +task manager, todo list, productivity app, task tracking +``` + +### App Store Connect Metadata Fields + +| Category | Field | Required | +|----------|-------|----------| +| **App Information** | Name | Yes | +| | Subtitle | No | +| | Category | Yes | +| | Secondary Category | No | +| | Content Rights | Yes | +| | Age Rating | Yes | +| **Version Information** | Description | Yes | +| | Keywords | Yes | +| | Promotional Text | No | +| | What's New | Yes (for updates) | +| | Support URL | Yes | +| | Marketing URL | No | +| **Pricing** | Price Tier | Yes | +| | Availability | Yes | + +### Age Rating Content Descriptors + +| Content Type | None | Infrequent | Frequent | +|--------------|------|------------|----------| +| Cartoon Violence | 4+ | 9+ | 12+ | +| Realistic Violence | 9+ | 12+ | 17+ | +| Sexual Content | 12+ | 17+ | 17+ | +| Profanity | 4+ | 12+ | 17+ | +| Alcohol/Drug Reference | 12+ | 17+ | 17+ | +| Gambling | 12+ | 17+ | 17+ | +| Horror/Fear | 9+ | 12+ | 17+ | + +--- + +## Google Play Store Requirements + +### Metadata Character Limits + +| Field | Character Limit | Notes | +|-------|----------------|-------| +| App Title | 50 characters | Increased from 30 in 2021 | +| Short Description | 80 characters | Visible on store listing | +| Full Description | 4,000 characters | Indexed for search keywords | +| Developer Name | 64 characters | Organization or individual | +| Developer Email | Required | Public support contact | +| Privacy Policy URL | Required | Must be valid HTTPS URL | + +### Description Keyword Strategy + +Google Play has no separate keyword field. Keywords are extracted from: + +1. **App Title** - Highest weight, most important +2. **Short Description** - High weight, visible in search +3. **Full Description** - Medium weight, use naturally throughout +4. **Developer Name** - Low weight but indexed + +**Keyword Density Guidelines:** +- Primary keyword: 2-3% density in full description +- Secondary keywords: 1-2% each +- Avoid keyword stuffing (>5% triggers spam detection) + +### Google Play Console Metadata + +| Category | Field | Required | +|----------|-------|----------| +| **Store Listing** | Title | Yes | +| | Short Description | Yes | +| | Full Description | Yes | +| | App Icon | Yes | +| | Feature Graphic | Yes | +| | Screenshots | Yes (min 2) | +| | Video | No | +| **Store Settings** | App Category | Yes | +| | Tags | No | +| | Contact Email | Yes | +| | Privacy Policy | Yes | +| **Content Rating** | IARC Questionnaire | Yes | + +### Content Rating (IARC) + +| Rating | Age | Description | +|--------|-----|-------------| +| PEGI 3 / Everyone | 3+ | Suitable for all ages | +| PEGI 7 / Everyone 10+ | 7+ | Mild violence, comic mischief | +| PEGI 12 / Teen | 12+ | Moderate violence, mild language | +| PEGI 16 / Mature 17+ | 16+ | Intense violence, strong language | +| PEGI 18 / Adults Only | 18+ | Extreme content | + +--- + +## Visual Asset Specifications + +### App Icon Requirements + +**Apple App Store:** + +| Device | Size | Format | +|--------|------|--------| +| iPhone | 1024x1024 px | PNG, no alpha | +| iPad | 1024x1024 px | PNG, no alpha | +| App Store | 1024x1024 px | PNG, no alpha | +| Spotlight | 120x120 px | PNG | +| Settings | 87x87 px | PNG | + +**Google Play Store:** + +| Asset | Size | Format | +|-------|------|--------| +| App Icon | 512x512 px | PNG, 32-bit | +| Feature Graphic | 1024x500 px | PNG or JPG | +| Promo Graphic | 180x120 px | PNG or JPG | +| TV Banner | 1280x720 px | PNG or JPG | + +### Screenshot Requirements + +**Apple App Store:** + +| Device | Portrait | Landscape | +|--------|----------|-----------| +| iPhone 6.9" | 1320x2868 px | 2868x1320 px | +| iPhone 6.5" | 1290x2796 px | 2796x1290 px | +| iPhone 5.5" | 1242x2208 px | 2208x1242 px | +| iPad Pro 12.9" | 2048x2732 px | 2732x2048 px | +| iPad 10.5" | 1668x2224 px | 2224x1668 px | + +- Minimum: 2 screenshots per device +- Maximum: 10 screenshots per device +- Format: PNG or JPG, no alpha channel +- First 3 screenshots are critical (most users don't scroll) + +**Google Play Store:** + +| Device | Dimensions | Notes | +|--------|------------|-------| +| Phone | 320-3840 px | Min 2:1 aspect ratio | +| 7" Tablet | 320-3840 px | Min 2:1 aspect ratio | +| 10" Tablet | 320-3840 px | Min 2:1 aspect ratio | +| Chromebook | 320-3840 px | Optional | +| TV | 320-3840 px | For TV apps only | + +- Minimum: 2 screenshots +- Maximum: 8 screenshots +- Format: PNG or JPG +- No transparency or borders + +### App Preview Video + +**Apple App Store:** +- Duration: 15-30 seconds +- Resolution: Match device screenshot size +- Format: M4V, MP4, MOV +- Frame rate: 30 fps +- Audio: Optional but recommended + +**Google Play Store:** +- YouTube video link only +- No duration limit (recommend under 2 minutes) +- Landscape orientation preferred +- Must not contain age-restricted content + +--- + +## Localization Requirements + +### Priority Markets by Revenue + +| Rank | Market | Language Code | +|------|--------|---------------| +| 1 | United States | en-US | +| 2 | Japan | ja | +| 3 | United Kingdom | en-GB | +| 4 | Germany | de-DE | +| 5 | China | zh-Hans (iOS), zh-CN (Android) | +| 6 | South Korea | ko | +| 7 | France | fr-FR | +| 8 | Canada | en-CA, fr-CA | +| 9 | Australia | en-AU | +| 10 | Russia | ru | + +### Apple App Store Localization + +Supported localizations: 40+ languages + +| Language | Locale Code | +|----------|-------------| +| English (US) | en-US | +| English (UK) | en-GB | +| Spanish | es-ES | +| Spanish (Mexico) | es-MX | +| French | fr-FR | +| German | de-DE | +| Japanese | ja | +| Korean | ko | +| Simplified Chinese | zh-Hans | +| Traditional Chinese | zh-Hant | + +### Google Play Store Localization + +Supported localizations: 75+ languages + +Each locale requires: +- Title (50 chars) +- Short description (80 chars) +- Full description (4,000 chars) +- Screenshots (can reuse or localize) + +--- + +## Compliance Guidelines + +### Apple App Store Review Guidelines Summary + +| Category | Key Requirements | +|----------|------------------| +| **Safety** | No objectionable content, privacy protection | +| **Performance** | App must work as described, no crashes | +| **Business** | Accurate app description, clear pricing | +| **Design** | Follow Human Interface Guidelines | +| **Legal** | Comply with local laws, proper licensing | + +**Common Rejection Reasons:** +1. Bugs and crashes (50%+ of rejections) +2. Broken links or placeholder content +3. Misleading app descriptions +4. Privacy policy missing or incomplete +5. In-app purchase issues + +### Google Play Developer Policies + +| Policy Area | Requirements | +|-------------|--------------| +| **Restricted Content** | No hate speech, violence, gambling (without license) | +| **Privacy** | Data collection disclosure, privacy policy | +| **Monetization** | Clear pricing, compliant IAPs | +| **Ads** | No deceptive ads, proper disclosure | +| **Store Listing** | Accurate description, no keyword stuffing | + +**Common Suspension Reasons:** +1. Policy violation (content, ads, permissions) +2. Repetitive content (clone apps) +3. Impersonation (fake apps) +4. Intellectual property infringement +5. Malicious behavior + +### Privacy Requirements + +**Apple (App Tracking Transparency):** +- ATT prompt required for tracking +- Privacy nutrition labels mandatory +- Data collection disclosure required + +**Google (Data Safety):** +- Data safety section mandatory +- Data collection and sharing disclosure +- Security practices declaration + +--- + +## Quick Reference Card + +### Apple vs Google Comparison + +| Attribute | Apple App Store | Google Play Store | +|-----------|-----------------|-------------------| +| Title Length | 30 chars | 50 chars | +| Subtitle | 30 chars | N/A | +| Short Description | N/A | 80 chars | +| Full Description | 4,000 chars | 4,000 chars | +| Keywords Field | 100 chars | N/A (in description) | +| Promotional Text | 170 chars | N/A | +| Icon Size | 1024x1024 px | 512x512 px | +| Min Screenshots | 2 | 2 | +| Max Screenshots | 10 | 8 | +| Review Time | 24-48 hours | 1-7 days | +| Metadata Update | Requires review | 1-2 hours to index | diff --git a/skills/app-store-optimization/sample_input.json b/skills/app-store-optimization/sample_input.json new file mode 100644 index 00000000..5435a367 --- /dev/null +++ b/skills/app-store-optimization/sample_input.json @@ -0,0 +1,30 @@ +{ + "request_type": "keyword_research", + "app_info": { + "name": "TaskFlow Pro", + "category": "Productivity", + "target_audience": "Professionals aged 25-45 working in teams", + "key_features": [ + "AI-powered task prioritization", + "Team collaboration tools", + "Calendar integration", + "Cross-platform sync" + ], + "unique_value": "AI automatically prioritizes your tasks based on deadlines and importance" + }, + "target_keywords": [ + "task manager", + "productivity app", + "todo list", + "team collaboration", + "project management" + ], + "competitors": [ + "Todoist", + "Any.do", + "Microsoft To Do", + "Things 3" + ], + "platform": "both", + "language": "en-US" +} diff --git a/skills/app-store-optimization/scripts/ab_test_planner.py b/skills/app-store-optimization/scripts/ab_test_planner.py new file mode 100644 index 00000000..06a80161 --- /dev/null +++ b/skills/app-store-optimization/scripts/ab_test_planner.py @@ -0,0 +1,662 @@ +""" +A/B testing module for App Store Optimization. +Plans and tracks A/B tests for metadata and visual assets. +""" + +from typing import Dict, List, Any, Optional +import math + + +class ABTestPlanner: + """Plans and tracks A/B tests for ASO elements.""" + + # Minimum detectable effect sizes (conservative estimates) + MIN_EFFECT_SIZES = { + 'icon': 0.10, # 10% conversion improvement + 'screenshot': 0.08, # 8% conversion improvement + 'title': 0.05, # 5% conversion improvement + 'description': 0.03 # 3% conversion improvement + } + + # Statistical confidence levels + CONFIDENCE_LEVELS = { + 'high': 0.95, # 95% confidence + 'standard': 0.90, # 90% confidence + 'exploratory': 0.80 # 80% confidence + } + + def __init__(self): + """Initialize A/B test planner.""" + self.active_tests = [] + + def design_test( + self, + test_type: str, + variant_a: Dict[str, Any], + variant_b: Dict[str, Any], + hypothesis: str, + success_metric: str = 'conversion_rate' + ) -> Dict[str, Any]: + """ + Design an A/B test with hypothesis and variables. + + Args: + test_type: Type of test ('icon', 'screenshot', 'title', 'description') + variant_a: Control variant details + variant_b: Test variant details + hypothesis: Expected outcome hypothesis + success_metric: Metric to optimize + + Returns: + Test design with configuration + """ + test_design = { + 'test_id': self._generate_test_id(test_type), + 'test_type': test_type, + 'hypothesis': hypothesis, + 'variants': { + 'a': { + 'name': 'Control', + 'details': variant_a, + 'traffic_split': 0.5 + }, + 'b': { + 'name': 'Variation', + 'details': variant_b, + 'traffic_split': 0.5 + } + }, + 'success_metric': success_metric, + 'secondary_metrics': self._get_secondary_metrics(test_type), + 'minimum_effect_size': self.MIN_EFFECT_SIZES.get(test_type, 0.05), + 'recommended_confidence': 'standard', + 'best_practices': self._get_test_best_practices(test_type) + } + + self.active_tests.append(test_design) + return test_design + + def calculate_sample_size( + self, + baseline_conversion: float, + minimum_detectable_effect: float, + confidence_level: str = 'standard', + power: float = 0.80 + ) -> Dict[str, Any]: + """ + Calculate required sample size for statistical significance. + + Args: + baseline_conversion: Current conversion rate (0-1) + minimum_detectable_effect: Minimum effect size to detect (0-1) + confidence_level: 'high', 'standard', or 'exploratory' + power: Statistical power (typically 0.80 or 0.90) + + Returns: + Sample size calculation with duration estimates + """ + alpha = 1 - self.CONFIDENCE_LEVELS[confidence_level] + beta = 1 - power + + # Expected conversion for variant B + expected_conversion_b = baseline_conversion * (1 + minimum_detectable_effect) + + # Z-scores for alpha and beta + z_alpha = self._get_z_score(1 - alpha / 2) # Two-tailed test + z_beta = self._get_z_score(power) + + # Pooled standard deviation + p_pooled = (baseline_conversion + expected_conversion_b) / 2 + sd_pooled = math.sqrt(2 * p_pooled * (1 - p_pooled)) + + # Sample size per variant + n_per_variant = math.ceil( + ((z_alpha + z_beta) ** 2 * sd_pooled ** 2) / + ((expected_conversion_b - baseline_conversion) ** 2) + ) + + total_sample_size = n_per_variant * 2 + + # Estimate duration based on typical traffic + duration_estimates = self._estimate_test_duration( + total_sample_size, + baseline_conversion + ) + + return { + 'sample_size_per_variant': n_per_variant, + 'total_sample_size': total_sample_size, + 'baseline_conversion': baseline_conversion, + 'expected_conversion_improvement': minimum_detectable_effect, + 'expected_conversion_b': expected_conversion_b, + 'confidence_level': confidence_level, + 'statistical_power': power, + 'duration_estimates': duration_estimates, + 'recommendations': self._generate_sample_size_recommendations( + n_per_variant, + duration_estimates + ) + } + + def calculate_significance( + self, + variant_a_conversions: int, + variant_a_visitors: int, + variant_b_conversions: int, + variant_b_visitors: int + ) -> Dict[str, Any]: + """ + Calculate statistical significance of test results. + + Args: + variant_a_conversions: Conversions for control + variant_a_visitors: Visitors for control + variant_b_conversions: Conversions for variation + variant_b_visitors: Visitors for variation + + Returns: + Significance analysis with decision recommendation + """ + # Calculate conversion rates + rate_a = variant_a_conversions / variant_a_visitors if variant_a_visitors > 0 else 0 + rate_b = variant_b_conversions / variant_b_visitors if variant_b_visitors > 0 else 0 + + # Calculate improvement + if rate_a > 0: + relative_improvement = (rate_b - rate_a) / rate_a + else: + relative_improvement = 0 + + absolute_improvement = rate_b - rate_a + + # Calculate standard error + se_a = math.sqrt(rate_a * (1 - rate_a) / variant_a_visitors) if variant_a_visitors > 0 else 0 + se_b = math.sqrt(rate_b * (1 - rate_b) / variant_b_visitors) if variant_b_visitors > 0 else 0 + se_diff = math.sqrt(se_a**2 + se_b**2) + + # Calculate z-score + z_score = absolute_improvement / se_diff if se_diff > 0 else 0 + + # Calculate p-value (two-tailed) + p_value = 2 * (1 - self._standard_normal_cdf(abs(z_score))) + + # Determine significance + is_significant_95 = p_value < 0.05 + is_significant_90 = p_value < 0.10 + + # Generate decision + decision = self._generate_test_decision( + relative_improvement, + is_significant_95, + is_significant_90, + variant_a_visitors + variant_b_visitors + ) + + return { + 'variant_a': { + 'conversions': variant_a_conversions, + 'visitors': variant_a_visitors, + 'conversion_rate': round(rate_a, 4) + }, + 'variant_b': { + 'conversions': variant_b_conversions, + 'visitors': variant_b_visitors, + 'conversion_rate': round(rate_b, 4) + }, + 'improvement': { + 'absolute': round(absolute_improvement, 4), + 'relative_percentage': round(relative_improvement * 100, 2) + }, + 'statistical_analysis': { + 'z_score': round(z_score, 3), + 'p_value': round(p_value, 4), + 'is_significant_95': is_significant_95, + 'is_significant_90': is_significant_90, + 'confidence_level': '95%' if is_significant_95 else ('90%' if is_significant_90 else 'Not significant') + }, + 'decision': decision + } + + def track_test_results( + self, + test_id: str, + results_data: Dict[str, Any] + ) -> Dict[str, Any]: + """ + Track ongoing test results and provide recommendations. + + Args: + test_id: Test identifier + results_data: Current test results + + Returns: + Test tracking report with next steps + """ + # Find test + test = next((t for t in self.active_tests if t['test_id'] == test_id), None) + if not test: + return {'error': f'Test {test_id} not found'} + + # Calculate significance + significance = self.calculate_significance( + results_data['variant_a_conversions'], + results_data['variant_a_visitors'], + results_data['variant_b_conversions'], + results_data['variant_b_visitors'] + ) + + # Calculate test progress + total_visitors = results_data['variant_a_visitors'] + results_data['variant_b_visitors'] + required_sample = results_data.get('required_sample_size', 10000) + progress_percentage = min((total_visitors / required_sample) * 100, 100) + + # Generate recommendations + recommendations = self._generate_tracking_recommendations( + significance, + progress_percentage, + test['test_type'] + ) + + return { + 'test_id': test_id, + 'test_type': test['test_type'], + 'progress': { + 'total_visitors': total_visitors, + 'required_sample_size': required_sample, + 'progress_percentage': round(progress_percentage, 1), + 'is_complete': progress_percentage >= 100 + }, + 'current_results': significance, + 'recommendations': recommendations, + 'next_steps': self._determine_next_steps( + significance, + progress_percentage + ) + } + + def generate_test_report( + self, + test_id: str, + final_results: Dict[str, Any] + ) -> Dict[str, Any]: + """ + Generate final test report with insights and recommendations. + + Args: + test_id: Test identifier + final_results: Final test results + + Returns: + Comprehensive test report + """ + test = next((t for t in self.active_tests if t['test_id'] == test_id), None) + if not test: + return {'error': f'Test {test_id} not found'} + + significance = self.calculate_significance( + final_results['variant_a_conversions'], + final_results['variant_a_visitors'], + final_results['variant_b_conversions'], + final_results['variant_b_visitors'] + ) + + # Generate insights + insights = self._generate_test_insights( + test, + significance, + final_results + ) + + # Implementation plan + implementation_plan = self._create_implementation_plan( + test, + significance + ) + + return { + 'test_summary': { + 'test_id': test_id, + 'test_type': test['test_type'], + 'hypothesis': test['hypothesis'], + 'duration_days': final_results.get('duration_days', 'N/A') + }, + 'results': significance, + 'insights': insights, + 'implementation_plan': implementation_plan, + 'learnings': self._extract_learnings(test, significance) + } + + def _generate_test_id(self, test_type: str) -> str: + """Generate unique test ID.""" + import time + timestamp = int(time.time()) + return f"{test_type}_{timestamp}" + + def _get_secondary_metrics(self, test_type: str) -> List[str]: + """Get secondary metrics to track for test type.""" + metrics_map = { + 'icon': ['tap_through_rate', 'impression_count', 'brand_recall'], + 'screenshot': ['tap_through_rate', 'time_on_page', 'scroll_depth'], + 'title': ['impression_count', 'tap_through_rate', 'search_visibility'], + 'description': ['time_on_page', 'scroll_depth', 'tap_through_rate'] + } + return metrics_map.get(test_type, ['tap_through_rate']) + + def _get_test_best_practices(self, test_type: str) -> List[str]: + """Get best practices for specific test type.""" + practices_map = { + 'icon': [ + 'Test only one element at a time (color vs. style vs. symbolism)', + 'Ensure icon is recognizable at small sizes (60x60px)', + 'Consider cultural context for global audience', + 'Test against top competitor icons' + ], + 'screenshot': [ + 'Test order of screenshots (users see first 2-3)', + 'Use captions to tell story', + 'Show key features and benefits', + 'Test with and without device frames' + ], + 'title': [ + 'Test keyword variations, not major rebrand', + 'Keep brand name consistent', + 'Ensure title fits within character limits', + 'Test on both search and browse contexts' + ], + 'description': [ + 'Test structure (bullet points vs. paragraphs)', + 'Test call-to-action placement', + 'Test feature vs. benefit focus', + 'Maintain keyword density' + ] + } + return practices_map.get(test_type, ['Test one variable at a time']) + + def _estimate_test_duration( + self, + required_sample_size: int, + baseline_conversion: float + ) -> Dict[str, Any]: + """Estimate test duration based on typical traffic levels.""" + # Assume different daily traffic scenarios + traffic_scenarios = { + 'low': 100, # 100 page views/day + 'medium': 1000, # 1000 page views/day + 'high': 10000 # 10000 page views/day + } + + estimates = {} + for scenario, daily_views in traffic_scenarios.items(): + days = math.ceil(required_sample_size / daily_views) + estimates[scenario] = { + 'daily_page_views': daily_views, + 'estimated_days': days, + 'estimated_weeks': round(days / 7, 1) + } + + return estimates + + def _generate_sample_size_recommendations( + self, + sample_size: int, + duration_estimates: Dict[str, Any] + ) -> List[str]: + """Generate recommendations based on sample size.""" + recommendations = [] + + if sample_size > 50000: + recommendations.append( + "Large sample size required - consider testing smaller effect size or increasing traffic" + ) + + if duration_estimates['medium']['estimated_days'] > 30: + recommendations.append( + "Long test duration - consider higher minimum detectable effect or focus on high-impact changes" + ) + + if duration_estimates['low']['estimated_days'] > 60: + recommendations.append( + "Insufficient traffic for reliable testing - consider user acquisition or broader targeting" + ) + + if not recommendations: + recommendations.append("Sample size and duration are reasonable for this test") + + return recommendations + + def _get_z_score(self, percentile: float) -> float: + """Get z-score for given percentile (approximation).""" + # Common z-scores + z_scores = { + 0.80: 0.84, + 0.85: 1.04, + 0.90: 1.28, + 0.95: 1.645, + 0.975: 1.96, + 0.99: 2.33 + } + return z_scores.get(percentile, 1.96) + + def _standard_normal_cdf(self, z: float) -> float: + """Approximate standard normal cumulative distribution function.""" + # Using error function approximation + t = 1.0 / (1.0 + 0.2316419 * abs(z)) + d = 0.3989423 * math.exp(-z * z / 2.0) + p = d * t * (0.3193815 + t * (-0.3565638 + t * (1.781478 + t * (-1.821256 + t * 1.330274)))) + + if z > 0: + return 1.0 - p + else: + return p + + def _generate_test_decision( + self, + improvement: float, + is_significant_95: bool, + is_significant_90: bool, + total_visitors: int + ) -> Dict[str, Any]: + """Generate test decision and recommendation.""" + if total_visitors < 1000: + return { + 'decision': 'continue', + 'rationale': 'Insufficient data - continue test to reach minimum sample size', + 'action': 'Keep test running' + } + + if is_significant_95: + if improvement > 0: + return { + 'decision': 'implement_b', + 'rationale': f'Variant B shows {improvement*100:.1f}% improvement with 95% confidence', + 'action': 'Implement Variant B' + } + else: + return { + 'decision': 'keep_a', + 'rationale': 'Variant A performs better with 95% confidence', + 'action': 'Keep current version (A)' + } + + elif is_significant_90: + if improvement > 0: + return { + 'decision': 'implement_b_cautiously', + 'rationale': f'Variant B shows {improvement*100:.1f}% improvement with 90% confidence', + 'action': 'Consider implementing B, monitor closely' + } + else: + return { + 'decision': 'keep_a', + 'rationale': 'Variant A performs better with 90% confidence', + 'action': 'Keep current version (A)' + } + + else: + return { + 'decision': 'inconclusive', + 'rationale': 'No statistically significant difference detected', + 'action': 'Either keep A or test different hypothesis' + } + + def _generate_tracking_recommendations( + self, + significance: Dict[str, Any], + progress: float, + test_type: str + ) -> List[str]: + """Generate recommendations for ongoing test.""" + recommendations = [] + + if progress < 50: + recommendations.append( + f"Test is {progress:.0f}% complete - continue collecting data" + ) + + if progress >= 100: + if significance['statistical_analysis']['is_significant_95']: + recommendations.append( + "Sufficient data collected with significant results - ready to conclude test" + ) + else: + recommendations.append( + "Sample size reached but no significant difference - consider extending test or concluding" + ) + + return recommendations + + def _determine_next_steps( + self, + significance: Dict[str, Any], + progress: float + ) -> str: + """Determine next steps for test.""" + if progress < 100: + return f"Continue test until reaching 100% sample size (currently {progress:.0f}%)" + + decision = significance.get('decision', {}).get('decision', 'inconclusive') + + if decision == 'implement_b': + return "Implement Variant B and monitor metrics for 2 weeks" + elif decision == 'keep_a': + return "Keep Variant A and design new test with different hypothesis" + else: + return "Test inconclusive - either keep A or design new test" + + def _generate_test_insights( + self, + test: Dict[str, Any], + significance: Dict[str, Any], + results: Dict[str, Any] + ) -> List[str]: + """Generate insights from test results.""" + insights = [] + + improvement = significance['improvement']['relative_percentage'] + + if significance['statistical_analysis']['is_significant_95']: + insights.append( + f"Strong evidence: Variant B {'improved' if improvement > 0 else 'decreased'} " + f"conversion by {abs(improvement):.1f}% with 95% confidence" + ) + + insights.append( + f"Tested {test['test_type']} changes: {test['hypothesis']}" + ) + + # Add context-specific insights + if test['test_type'] == 'icon' and improvement > 5: + insights.append( + "Icon change had substantial impact - visual first impression is critical" + ) + + return insights + + def _create_implementation_plan( + self, + test: Dict[str, Any], + significance: Dict[str, Any] + ) -> List[Dict[str, str]]: + """Create implementation plan for winning variant.""" + plan = [] + + if significance.get('decision', {}).get('decision') == 'implement_b': + plan.append({ + 'step': '1. Update store listing', + 'details': f"Replace {test['test_type']} with Variant B across all platforms" + }) + plan.append({ + 'step': '2. Monitor metrics', + 'details': 'Track conversion rate for 2 weeks to confirm sustained improvement' + }) + plan.append({ + 'step': '3. Document learnings', + 'details': 'Record insights for future optimization' + }) + + return plan + + def _extract_learnings( + self, + test: Dict[str, Any], + significance: Dict[str, Any] + ) -> List[str]: + """Extract key learnings from test.""" + learnings = [] + + improvement = significance['improvement']['relative_percentage'] + + learnings.append( + f"Testing {test['test_type']} can yield {abs(improvement):.1f}% conversion change" + ) + + if test['test_type'] == 'title': + learnings.append( + "Title changes affect search visibility and user perception" + ) + elif test['test_type'] == 'screenshot': + learnings.append( + "First 2-3 screenshots are critical for conversion" + ) + + return learnings + + +def plan_ab_test( + test_type: str, + variant_a: Dict[str, Any], + variant_b: Dict[str, Any], + hypothesis: str, + baseline_conversion: float +) -> Dict[str, Any]: + """ + Convenience function to plan an A/B test. + + Args: + test_type: Type of test + variant_a: Control variant + variant_b: Test variant + hypothesis: Test hypothesis + baseline_conversion: Current conversion rate + + Returns: + Complete test plan + """ + planner = ABTestPlanner() + + test_design = planner.design_test( + test_type, + variant_a, + variant_b, + hypothesis + ) + + sample_size = planner.calculate_sample_size( + baseline_conversion, + planner.MIN_EFFECT_SIZES.get(test_type, 0.05) + ) + + return { + 'test_design': test_design, + 'sample_size_requirements': sample_size + } diff --git a/skills/app-store-optimization/scripts/aso_scorer.py b/skills/app-store-optimization/scripts/aso_scorer.py new file mode 100644 index 00000000..ba4ea6ac --- /dev/null +++ b/skills/app-store-optimization/scripts/aso_scorer.py @@ -0,0 +1,482 @@ +""" +ASO scoring module for App Store Optimization. +Calculates comprehensive ASO health score across multiple dimensions. +""" + +from typing import Dict, List, Any, Optional + + +class ASOScorer: + """Calculates overall ASO health score and provides recommendations.""" + + # Score weights for different components (total = 100) + WEIGHTS = { + 'metadata_quality': 25, + 'ratings_reviews': 25, + 'keyword_performance': 25, + 'conversion_metrics': 25 + } + + # Benchmarks for scoring + BENCHMARKS = { + 'title_keyword_usage': {'min': 1, 'target': 2}, + 'description_length': {'min': 500, 'target': 2000}, + 'keyword_density': {'min': 2, 'optimal': 5, 'max': 8}, + 'average_rating': {'min': 3.5, 'target': 4.5}, + 'ratings_count': {'min': 100, 'target': 5000}, + 'keywords_top_10': {'min': 2, 'target': 10}, + 'keywords_top_50': {'min': 5, 'target': 20}, + 'conversion_rate': {'min': 0.02, 'target': 0.10} + } + + def __init__(self): + """Initialize ASO scorer.""" + self.score_breakdown = {} + + def calculate_overall_score( + self, + metadata: Dict[str, Any], + ratings: Dict[str, Any], + keyword_performance: Dict[str, Any], + conversion: Dict[str, Any] + ) -> Dict[str, Any]: + """ + Calculate comprehensive ASO score (0-100). + + Args: + metadata: Title, description quality metrics + ratings: Rating average and count + keyword_performance: Keyword ranking data + conversion: Impression-to-install metrics + + Returns: + Overall score with detailed breakdown + """ + # Calculate component scores + metadata_score = self.score_metadata_quality(metadata) + ratings_score = self.score_ratings_reviews(ratings) + keyword_score = self.score_keyword_performance(keyword_performance) + conversion_score = self.score_conversion_metrics(conversion) + + # Calculate weighted overall score + overall_score = ( + metadata_score * (self.WEIGHTS['metadata_quality'] / 100) + + ratings_score * (self.WEIGHTS['ratings_reviews'] / 100) + + keyword_score * (self.WEIGHTS['keyword_performance'] / 100) + + conversion_score * (self.WEIGHTS['conversion_metrics'] / 100) + ) + + # Store breakdown + self.score_breakdown = { + 'metadata_quality': { + 'score': metadata_score, + 'weight': self.WEIGHTS['metadata_quality'], + 'weighted_contribution': round(metadata_score * (self.WEIGHTS['metadata_quality'] / 100), 1) + }, + 'ratings_reviews': { + 'score': ratings_score, + 'weight': self.WEIGHTS['ratings_reviews'], + 'weighted_contribution': round(ratings_score * (self.WEIGHTS['ratings_reviews'] / 100), 1) + }, + 'keyword_performance': { + 'score': keyword_score, + 'weight': self.WEIGHTS['keyword_performance'], + 'weighted_contribution': round(keyword_score * (self.WEIGHTS['keyword_performance'] / 100), 1) + }, + 'conversion_metrics': { + 'score': conversion_score, + 'weight': self.WEIGHTS['conversion_metrics'], + 'weighted_contribution': round(conversion_score * (self.WEIGHTS['conversion_metrics'] / 100), 1) + } + } + + # Generate recommendations + recommendations = self.generate_recommendations( + metadata_score, + ratings_score, + keyword_score, + conversion_score + ) + + # Assess overall health + health_status = self._assess_health_status(overall_score) + + return { + 'overall_score': round(overall_score, 1), + 'health_status': health_status, + 'score_breakdown': self.score_breakdown, + 'recommendations': recommendations, + 'priority_actions': self._prioritize_actions(recommendations), + 'strengths': self._identify_strengths(self.score_breakdown), + 'weaknesses': self._identify_weaknesses(self.score_breakdown) + } + + def score_metadata_quality(self, metadata: Dict[str, Any]) -> float: + """ + Score metadata quality (0-100). + + Evaluates: + - Title optimization + - Description quality + - Keyword usage + """ + scores = [] + + # Title score (0-35 points) + title_keywords = metadata.get('title_keyword_count', 0) + title_length = metadata.get('title_length', 0) + + title_score = 0 + if title_keywords >= self.BENCHMARKS['title_keyword_usage']['target']: + title_score = 35 + elif title_keywords >= self.BENCHMARKS['title_keyword_usage']['min']: + title_score = 25 + else: + title_score = 10 + + # Adjust for title length usage + if title_length > 25: # Using most of available space + title_score += 0 + else: + title_score -= 5 + + scores.append(min(title_score, 35)) + + # Description score (0-35 points) + desc_length = metadata.get('description_length', 0) + desc_quality = metadata.get('description_quality', 0.0) # 0-1 scale + + desc_score = 0 + if desc_length >= self.BENCHMARKS['description_length']['target']: + desc_score = 25 + elif desc_length >= self.BENCHMARKS['description_length']['min']: + desc_score = 15 + else: + desc_score = 5 + + # Add quality bonus + desc_score += desc_quality * 10 + scores.append(min(desc_score, 35)) + + # Keyword density score (0-30 points) + keyword_density = metadata.get('keyword_density', 0.0) + + if self.BENCHMARKS['keyword_density']['min'] <= keyword_density <= self.BENCHMARKS['keyword_density']['optimal']: + density_score = 30 + elif keyword_density < self.BENCHMARKS['keyword_density']['min']: + # Too low - proportional scoring + density_score = (keyword_density / self.BENCHMARKS['keyword_density']['min']) * 20 + else: + # Too high (keyword stuffing) - penalty + excess = keyword_density - self.BENCHMARKS['keyword_density']['optimal'] + density_score = max(30 - (excess * 5), 0) + + scores.append(density_score) + + return round(sum(scores), 1) + + def score_ratings_reviews(self, ratings: Dict[str, Any]) -> float: + """ + Score ratings and reviews (0-100). + + Evaluates: + - Average rating + - Total ratings count + - Review velocity + """ + average_rating = ratings.get('average_rating', 0.0) + total_ratings = ratings.get('total_ratings', 0) + recent_ratings = ratings.get('recent_ratings_30d', 0) + + # Rating quality score (0-50 points) + if average_rating >= self.BENCHMARKS['average_rating']['target']: + rating_quality_score = 50 + elif average_rating >= self.BENCHMARKS['average_rating']['min']: + # Proportional scoring between min and target + proportion = (average_rating - self.BENCHMARKS['average_rating']['min']) / \ + (self.BENCHMARKS['average_rating']['target'] - self.BENCHMARKS['average_rating']['min']) + rating_quality_score = 30 + (proportion * 20) + elif average_rating >= 3.0: + rating_quality_score = 20 + else: + rating_quality_score = 10 + + # Rating volume score (0-30 points) + if total_ratings >= self.BENCHMARKS['ratings_count']['target']: + rating_volume_score = 30 + elif total_ratings >= self.BENCHMARKS['ratings_count']['min']: + # Proportional scoring + proportion = (total_ratings - self.BENCHMARKS['ratings_count']['min']) / \ + (self.BENCHMARKS['ratings_count']['target'] - self.BENCHMARKS['ratings_count']['min']) + rating_volume_score = 15 + (proportion * 15) + else: + # Very low volume + rating_volume_score = (total_ratings / self.BENCHMARKS['ratings_count']['min']) * 15 + + # Rating velocity score (0-20 points) + if recent_ratings > 100: + velocity_score = 20 + elif recent_ratings > 50: + velocity_score = 15 + elif recent_ratings > 10: + velocity_score = 10 + else: + velocity_score = 5 + + total_score = rating_quality_score + rating_volume_score + velocity_score + + return round(min(total_score, 100), 1) + + def score_keyword_performance(self, keyword_performance: Dict[str, Any]) -> float: + """ + Score keyword ranking performance (0-100). + + Evaluates: + - Top 10 rankings + - Top 50 rankings + - Ranking trends + """ + top_10_count = keyword_performance.get('top_10', 0) + top_50_count = keyword_performance.get('top_50', 0) + top_100_count = keyword_performance.get('top_100', 0) + improving_keywords = keyword_performance.get('improving_keywords', 0) + + # Top 10 score (0-50 points) - most valuable rankings + if top_10_count >= self.BENCHMARKS['keywords_top_10']['target']: + top_10_score = 50 + elif top_10_count >= self.BENCHMARKS['keywords_top_10']['min']: + proportion = (top_10_count - self.BENCHMARKS['keywords_top_10']['min']) / \ + (self.BENCHMARKS['keywords_top_10']['target'] - self.BENCHMARKS['keywords_top_10']['min']) + top_10_score = 25 + (proportion * 25) + else: + top_10_score = (top_10_count / self.BENCHMARKS['keywords_top_10']['min']) * 25 + + # Top 50 score (0-30 points) + if top_50_count >= self.BENCHMARKS['keywords_top_50']['target']: + top_50_score = 30 + elif top_50_count >= self.BENCHMARKS['keywords_top_50']['min']: + proportion = (top_50_count - self.BENCHMARKS['keywords_top_50']['min']) / \ + (self.BENCHMARKS['keywords_top_50']['target'] - self.BENCHMARKS['keywords_top_50']['min']) + top_50_score = 15 + (proportion * 15) + else: + top_50_score = (top_50_count / self.BENCHMARKS['keywords_top_50']['min']) * 15 + + # Coverage score (0-10 points) - based on top 100 + coverage_score = min((top_100_count / 30) * 10, 10) + + # Trend score (0-10 points) - are rankings improving? + if improving_keywords > 5: + trend_score = 10 + elif improving_keywords > 0: + trend_score = 5 + else: + trend_score = 0 + + total_score = top_10_score + top_50_score + coverage_score + trend_score + + return round(min(total_score, 100), 1) + + def score_conversion_metrics(self, conversion: Dict[str, Any]) -> float: + """ + Score conversion performance (0-100). + + Evaluates: + - Impression-to-install conversion rate + - Download velocity + """ + conversion_rate = conversion.get('impression_to_install', 0.0) + downloads_30d = conversion.get('downloads_last_30_days', 0) + downloads_trend = conversion.get('downloads_trend', 'stable') # 'up', 'stable', 'down' + + # Conversion rate score (0-70 points) + if conversion_rate >= self.BENCHMARKS['conversion_rate']['target']: + conversion_score = 70 + elif conversion_rate >= self.BENCHMARKS['conversion_rate']['min']: + proportion = (conversion_rate - self.BENCHMARKS['conversion_rate']['min']) / \ + (self.BENCHMARKS['conversion_rate']['target'] - self.BENCHMARKS['conversion_rate']['min']) + conversion_score = 35 + (proportion * 35) + else: + conversion_score = (conversion_rate / self.BENCHMARKS['conversion_rate']['min']) * 35 + + # Download velocity score (0-20 points) + if downloads_30d > 10000: + velocity_score = 20 + elif downloads_30d > 1000: + velocity_score = 15 + elif downloads_30d > 100: + velocity_score = 10 + else: + velocity_score = 5 + + # Trend bonus (0-10 points) + if downloads_trend == 'up': + trend_score = 10 + elif downloads_trend == 'stable': + trend_score = 5 + else: + trend_score = 0 + + total_score = conversion_score + velocity_score + trend_score + + return round(min(total_score, 100), 1) + + def generate_recommendations( + self, + metadata_score: float, + ratings_score: float, + keyword_score: float, + conversion_score: float + ) -> List[Dict[str, Any]]: + """Generate prioritized recommendations based on scores.""" + recommendations = [] + + # Metadata recommendations + if metadata_score < 60: + recommendations.append({ + 'category': 'metadata_quality', + 'priority': 'high', + 'action': 'Optimize app title and description', + 'details': 'Add more keywords to title, expand description to 1500-2000 characters, improve keyword density to 3-5%', + 'expected_impact': 'Improve discoverability and ranking potential' + }) + elif metadata_score < 80: + recommendations.append({ + 'category': 'metadata_quality', + 'priority': 'medium', + 'action': 'Refine metadata for better keyword targeting', + 'details': 'Test variations of title/subtitle, optimize keyword field for Apple', + 'expected_impact': 'Incremental ranking improvements' + }) + + # Ratings recommendations + if ratings_score < 60: + recommendations.append({ + 'category': 'ratings_reviews', + 'priority': 'high', + 'action': 'Improve rating quality and volume', + 'details': 'Address top user complaints, implement in-app rating prompts, respond to negative reviews', + 'expected_impact': 'Better conversion rates and trust signals' + }) + elif ratings_score < 80: + recommendations.append({ + 'category': 'ratings_reviews', + 'priority': 'medium', + 'action': 'Increase rating velocity', + 'details': 'Optimize timing of rating requests, encourage satisfied users to rate', + 'expected_impact': 'Sustained rating quality' + }) + + # Keyword performance recommendations + if keyword_score < 60: + recommendations.append({ + 'category': 'keyword_performance', + 'priority': 'high', + 'action': 'Improve keyword rankings', + 'details': 'Target long-tail keywords with lower competition, update metadata with high-potential keywords, build backlinks', + 'expected_impact': 'Significant improvement in organic visibility' + }) + elif keyword_score < 80: + recommendations.append({ + 'category': 'keyword_performance', + 'priority': 'medium', + 'action': 'Expand keyword coverage', + 'details': 'Target additional related keywords, test seasonal keywords, localize for new markets', + 'expected_impact': 'Broader reach and more discovery opportunities' + }) + + # Conversion recommendations + if conversion_score < 60: + recommendations.append({ + 'category': 'conversion_metrics', + 'priority': 'high', + 'action': 'Optimize store listing for conversions', + 'details': 'Improve screenshots and icon, strengthen value proposition in description, add video preview', + 'expected_impact': 'Higher impression-to-install conversion' + }) + elif conversion_score < 80: + recommendations.append({ + 'category': 'conversion_metrics', + 'priority': 'medium', + 'action': 'Test visual asset variations', + 'details': 'A/B test different icon designs and screenshot sequences', + 'expected_impact': 'Incremental conversion improvements' + }) + + return recommendations + + def _assess_health_status(self, overall_score: float) -> str: + """Assess overall ASO health status.""" + if overall_score >= 80: + return "Excellent - Top-tier ASO performance" + elif overall_score >= 65: + return "Good - Competitive ASO with room for improvement" + elif overall_score >= 50: + return "Fair - Needs strategic improvements" + else: + return "Poor - Requires immediate ASO overhaul" + + def _prioritize_actions( + self, + recommendations: List[Dict[str, Any]] + ) -> List[Dict[str, Any]]: + """Prioritize actions by impact and urgency.""" + # Sort by priority (high first) and expected impact + priority_order = {'high': 0, 'medium': 1, 'low': 2} + + sorted_recommendations = sorted( + recommendations, + key=lambda x: priority_order[x['priority']] + ) + + return sorted_recommendations[:3] # Top 3 priority actions + + def _identify_strengths(self, score_breakdown: Dict[str, Any]) -> List[str]: + """Identify areas of strength (scores >= 75).""" + strengths = [] + + for category, data in score_breakdown.items(): + if data['score'] >= 75: + strengths.append( + f"{category.replace('_', ' ').title()}: {data['score']}/100" + ) + + return strengths if strengths else ["Focus on building strengths across all areas"] + + def _identify_weaknesses(self, score_breakdown: Dict[str, Any]) -> List[str]: + """Identify areas needing improvement (scores < 60).""" + weaknesses = [] + + for category, data in score_breakdown.items(): + if data['score'] < 60: + weaknesses.append( + f"{category.replace('_', ' ').title()}: {data['score']}/100 - needs improvement" + ) + + return weaknesses if weaknesses else ["All areas performing adequately"] + + +def calculate_aso_score( + metadata: Dict[str, Any], + ratings: Dict[str, Any], + keyword_performance: Dict[str, Any], + conversion: Dict[str, Any] +) -> Dict[str, Any]: + """ + Convenience function to calculate ASO score. + + Args: + metadata: Metadata quality metrics + ratings: Ratings data + keyword_performance: Keyword ranking data + conversion: Conversion metrics + + Returns: + Complete ASO score report + """ + scorer = ASOScorer() + return scorer.calculate_overall_score( + metadata, + ratings, + keyword_performance, + conversion + ) diff --git a/skills/app-store-optimization/scripts/competitor_analyzer.py b/skills/app-store-optimization/scripts/competitor_analyzer.py new file mode 100644 index 00000000..9f84575b --- /dev/null +++ b/skills/app-store-optimization/scripts/competitor_analyzer.py @@ -0,0 +1,577 @@ +""" +Competitor analysis module for App Store Optimization. +Analyzes top competitors' ASO strategies and identifies opportunities. +""" + +from typing import Dict, List, Any, Optional +from collections import Counter +import re + + +class CompetitorAnalyzer: + """Analyzes competitor apps to identify ASO opportunities.""" + + def __init__(self, category: str, platform: str = 'apple'): + """ + Initialize competitor analyzer. + + Args: + category: App category (e.g., "Productivity", "Games") + platform: 'apple' or 'google' + """ + self.category = category + self.platform = platform + self.competitors = [] + + def analyze_competitor( + self, + app_data: Dict[str, Any] + ) -> Dict[str, Any]: + """ + Analyze a single competitor's ASO strategy. + + Args: + app_data: Dictionary with app_name, title, description, rating, ratings_count, keywords + + Returns: + Comprehensive competitor analysis + """ + app_name = app_data.get('app_name', '') + title = app_data.get('title', '') + description = app_data.get('description', '') + rating = app_data.get('rating', 0.0) + ratings_count = app_data.get('ratings_count', 0) + keywords = app_data.get('keywords', []) + + analysis = { + 'app_name': app_name, + 'title_analysis': self._analyze_title(title), + 'description_analysis': self._analyze_description(description), + 'keyword_strategy': self._extract_keyword_strategy(title, description, keywords), + 'rating_metrics': { + 'rating': rating, + 'ratings_count': ratings_count, + 'rating_quality': self._assess_rating_quality(rating, ratings_count) + }, + 'competitive_strength': self._calculate_competitive_strength( + rating, + ratings_count, + len(description) + ), + 'key_differentiators': self._identify_differentiators(description) + } + + self.competitors.append(analysis) + return analysis + + def compare_competitors( + self, + competitors_data: List[Dict[str, Any]] + ) -> Dict[str, Any]: + """ + Compare multiple competitors and identify patterns. + + Args: + competitors_data: List of competitor data dictionaries + + Returns: + Comparative analysis with insights + """ + # Analyze each competitor + analyses = [] + for comp_data in competitors_data: + analysis = self.analyze_competitor(comp_data) + analyses.append(analysis) + + # Extract common keywords across competitors + all_keywords = [] + for analysis in analyses: + all_keywords.extend(analysis['keyword_strategy']['primary_keywords']) + + common_keywords = self._find_common_keywords(all_keywords) + + # Identify keyword gaps (used by some but not all) + keyword_gaps = self._identify_keyword_gaps(analyses) + + # Rank competitors by strength + ranked_competitors = sorted( + analyses, + key=lambda x: x['competitive_strength'], + reverse=True + ) + + # Analyze rating distribution + rating_analysis = self._analyze_rating_distribution(analyses) + + # Identify best practices + best_practices = self._identify_best_practices(ranked_competitors) + + return { + 'category': self.category, + 'platform': self.platform, + 'competitors_analyzed': len(analyses), + 'ranked_competitors': ranked_competitors, + 'common_keywords': common_keywords, + 'keyword_gaps': keyword_gaps, + 'rating_analysis': rating_analysis, + 'best_practices': best_practices, + 'opportunities': self._identify_opportunities( + analyses, + common_keywords, + keyword_gaps + ) + } + + def identify_gaps( + self, + your_app_data: Dict[str, Any], + competitors_data: List[Dict[str, Any]] + ) -> Dict[str, Any]: + """ + Identify gaps between your app and competitors. + + Args: + your_app_data: Your app's data + competitors_data: List of competitor data + + Returns: + Gap analysis with actionable recommendations + """ + # Analyze your app + your_analysis = self.analyze_competitor(your_app_data) + + # Analyze competitors + competitor_comparison = self.compare_competitors(competitors_data) + + # Identify keyword gaps + your_keywords = set(your_analysis['keyword_strategy']['primary_keywords']) + competitor_keywords = set(competitor_comparison['common_keywords']) + missing_keywords = competitor_keywords - your_keywords + + # Identify rating gap + avg_competitor_rating = competitor_comparison['rating_analysis']['average_rating'] + rating_gap = avg_competitor_rating - your_analysis['rating_metrics']['rating'] + + # Identify description length gap + avg_competitor_desc_length = sum( + len(comp['description_analysis']['text']) + for comp in competitor_comparison['ranked_competitors'] + ) / len(competitor_comparison['ranked_competitors']) + your_desc_length = len(your_analysis['description_analysis']['text']) + desc_length_gap = avg_competitor_desc_length - your_desc_length + + return { + 'your_app': your_analysis, + 'keyword_gaps': { + 'missing_keywords': list(missing_keywords)[:10], + 'recommendations': self._generate_keyword_recommendations(missing_keywords) + }, + 'rating_gap': { + 'your_rating': your_analysis['rating_metrics']['rating'], + 'average_competitor_rating': avg_competitor_rating, + 'gap': round(rating_gap, 2), + 'action_items': self._generate_rating_improvement_actions(rating_gap) + }, + 'content_gap': { + 'your_description_length': your_desc_length, + 'average_competitor_length': int(avg_competitor_desc_length), + 'gap': int(desc_length_gap), + 'recommendations': self._generate_content_recommendations(desc_length_gap) + }, + 'competitive_positioning': self._assess_competitive_position( + your_analysis, + competitor_comparison + ) + } + + def _analyze_title(self, title: str) -> Dict[str, Any]: + """Analyze title structure and keyword usage.""" + parts = re.split(r'[-:|]', title) + + return { + 'title': title, + 'length': len(title), + 'has_brand': len(parts) > 0, + 'has_keywords': len(parts) > 1, + 'components': [part.strip() for part in parts], + 'word_count': len(title.split()), + 'strategy': 'brand_plus_keywords' if len(parts) > 1 else 'brand_only' + } + + def _analyze_description(self, description: str) -> Dict[str, Any]: + """Analyze description structure and content.""" + lines = description.split('\n') + word_count = len(description.split()) + + # Check for structural elements + has_bullet_points = '•' in description or '*' in description + has_sections = any(line.isupper() for line in lines if len(line) > 0) + has_call_to_action = any( + cta in description.lower() + for cta in ['download', 'try', 'get', 'start', 'join'] + ) + + # Extract features mentioned + features = self._extract_features(description) + + return { + 'text': description, + 'length': len(description), + 'word_count': word_count, + 'structure': { + 'has_bullet_points': has_bullet_points, + 'has_sections': has_sections, + 'has_call_to_action': has_call_to_action + }, + 'features_mentioned': features, + 'readability': 'good' if 50 <= word_count <= 300 else 'needs_improvement' + } + + def _extract_keyword_strategy( + self, + title: str, + description: str, + explicit_keywords: List[str] + ) -> Dict[str, Any]: + """Extract keyword strategy from metadata.""" + # Extract keywords from title + title_keywords = [word.lower() for word in title.split() if len(word) > 3] + + # Extract frequently used words from description + desc_words = re.findall(r'\b\w{4,}\b', description.lower()) + word_freq = Counter(desc_words) + frequent_words = [word for word, count in word_freq.most_common(15) if count > 2] + + # Combine with explicit keywords + all_keywords = list(set(title_keywords + frequent_words + explicit_keywords)) + + return { + 'primary_keywords': title_keywords, + 'description_keywords': frequent_words[:10], + 'explicit_keywords': explicit_keywords, + 'total_unique_keywords': len(all_keywords), + 'keyword_focus': self._assess_keyword_focus(title_keywords, frequent_words) + } + + def _assess_rating_quality(self, rating: float, ratings_count: int) -> str: + """Assess the quality of ratings.""" + if ratings_count < 100: + return 'insufficient_data' + elif rating >= 4.5 and ratings_count > 1000: + return 'excellent' + elif rating >= 4.0 and ratings_count > 500: + return 'good' + elif rating >= 3.5: + return 'average' + else: + return 'poor' + + def _calculate_competitive_strength( + self, + rating: float, + ratings_count: int, + description_length: int + ) -> float: + """ + Calculate overall competitive strength (0-100). + + Factors: + - Rating quality (40%) + - Rating volume (30%) + - Metadata quality (30%) + """ + # Rating quality score (0-40) + rating_score = (rating / 5.0) * 40 + + # Rating volume score (0-30) + volume_score = min((ratings_count / 10000) * 30, 30) + + # Metadata quality score (0-30) + metadata_score = min((description_length / 2000) * 30, 30) + + total_score = rating_score + volume_score + metadata_score + + return round(total_score, 1) + + def _identify_differentiators(self, description: str) -> List[str]: + """Identify key differentiators from description.""" + differentiator_keywords = [ + 'unique', 'only', 'first', 'best', 'leading', 'exclusive', + 'revolutionary', 'innovative', 'patent', 'award' + ] + + differentiators = [] + sentences = description.split('.') + + for sentence in sentences: + sentence_lower = sentence.lower() + if any(keyword in sentence_lower for keyword in differentiator_keywords): + differentiators.append(sentence.strip()) + + return differentiators[:5] + + def _find_common_keywords(self, all_keywords: List[str]) -> List[str]: + """Find keywords used by multiple competitors.""" + keyword_counts = Counter(all_keywords) + # Return keywords used by at least 2 competitors + common = [kw for kw, count in keyword_counts.items() if count >= 2] + return sorted(common, key=lambda x: keyword_counts[x], reverse=True)[:20] + + def _identify_keyword_gaps(self, analyses: List[Dict[str, Any]]) -> List[Dict[str, Any]]: + """Identify keywords used by some competitors but not others.""" + all_keywords_by_app = {} + + for analysis in analyses: + app_name = analysis['app_name'] + keywords = analysis['keyword_strategy']['primary_keywords'] + all_keywords_by_app[app_name] = set(keywords) + + # Find keywords used by some but not all + all_keywords_set = set() + for keywords in all_keywords_by_app.values(): + all_keywords_set.update(keywords) + + gaps = [] + for keyword in all_keywords_set: + using_apps = [ + app for app, keywords in all_keywords_by_app.items() + if keyword in keywords + ] + if 1 < len(using_apps) < len(analyses): + gaps.append({ + 'keyword': keyword, + 'used_by': using_apps, + 'usage_percentage': round(len(using_apps) / len(analyses) * 100, 1) + }) + + return sorted(gaps, key=lambda x: x['usage_percentage'], reverse=True)[:15] + + def _analyze_rating_distribution(self, analyses: List[Dict[str, Any]]) -> Dict[str, Any]: + """Analyze rating distribution across competitors.""" + ratings = [a['rating_metrics']['rating'] for a in analyses] + ratings_counts = [a['rating_metrics']['ratings_count'] for a in analyses] + + return { + 'average_rating': round(sum(ratings) / len(ratings), 2), + 'highest_rating': max(ratings), + 'lowest_rating': min(ratings), + 'average_ratings_count': int(sum(ratings_counts) / len(ratings_counts)), + 'total_ratings_in_category': sum(ratings_counts) + } + + def _identify_best_practices(self, ranked_competitors: List[Dict[str, Any]]) -> List[str]: + """Identify best practices from top competitors.""" + if not ranked_competitors: + return [] + + top_competitor = ranked_competitors[0] + practices = [] + + # Title strategy + title_analysis = top_competitor['title_analysis'] + if title_analysis['has_keywords']: + practices.append( + f"Title Strategy: Include primary keyword in title (e.g., '{title_analysis['title']}')" + ) + + # Description structure + desc_analysis = top_competitor['description_analysis'] + if desc_analysis['structure']['has_bullet_points']: + practices.append("Description: Use bullet points to highlight key features") + + if desc_analysis['structure']['has_sections']: + practices.append("Description: Organize content with clear section headers") + + # Rating strategy + rating_quality = top_competitor['rating_metrics']['rating_quality'] + if rating_quality in ['excellent', 'good']: + practices.append( + f"Ratings: Maintain high rating quality ({top_competitor['rating_metrics']['rating']}★) " + f"with significant volume ({top_competitor['rating_metrics']['ratings_count']} ratings)" + ) + + return practices[:5] + + def _identify_opportunities( + self, + analyses: List[Dict[str, Any]], + common_keywords: List[str], + keyword_gaps: List[Dict[str, Any]] + ) -> List[str]: + """Identify ASO opportunities based on competitive analysis.""" + opportunities = [] + + # Keyword opportunities from gaps + if keyword_gaps: + underutilized_keywords = [ + gap['keyword'] for gap in keyword_gaps + if gap['usage_percentage'] < 50 + ] + if underutilized_keywords: + opportunities.append( + f"Target underutilized keywords: {', '.join(underutilized_keywords[:5])}" + ) + + # Rating opportunity + avg_rating = sum(a['rating_metrics']['rating'] for a in analyses) / len(analyses) + if avg_rating < 4.5: + opportunities.append( + f"Category average rating is {avg_rating:.1f} - opportunity to differentiate with higher ratings" + ) + + # Content depth opportunity + avg_desc_length = sum( + a['description_analysis']['length'] for a in analyses + ) / len(analyses) + if avg_desc_length < 1500: + opportunities.append( + "Competitors have relatively short descriptions - opportunity to provide more comprehensive information" + ) + + return opportunities[:5] + + def _extract_features(self, description: str) -> List[str]: + """Extract feature mentions from description.""" + # Look for bullet points or numbered lists + lines = description.split('\n') + features = [] + + for line in lines: + line = line.strip() + # Check if line starts with bullet or number + if line and (line[0] in ['•', '*', '-', '✓'] or line[0].isdigit()): + # Clean the line + cleaned = re.sub(r'^[•*\-✓\d.)\s]+', '', line) + if cleaned: + features.append(cleaned) + + return features[:10] + + def _assess_keyword_focus( + self, + title_keywords: List[str], + description_keywords: List[str] + ) -> str: + """Assess keyword focus strategy.""" + overlap = set(title_keywords) & set(description_keywords) + + if len(overlap) >= 3: + return 'consistent_focus' + elif len(overlap) >= 1: + return 'moderate_focus' + else: + return 'broad_focus' + + def _generate_keyword_recommendations(self, missing_keywords: set) -> List[str]: + """Generate recommendations for missing keywords.""" + if not missing_keywords: + return ["Your keyword coverage is comprehensive"] + + recommendations = [] + missing_list = list(missing_keywords)[:5] + + recommendations.append( + f"Consider adding these competitor keywords: {', '.join(missing_list)}" + ) + recommendations.append( + "Test keyword variations in subtitle/promotional text first" + ) + recommendations.append( + "Monitor competitor keyword changes monthly" + ) + + return recommendations + + def _generate_rating_improvement_actions(self, rating_gap: float) -> List[str]: + """Generate actions to improve ratings.""" + actions = [] + + if rating_gap > 0.5: + actions.append("CRITICAL: Significant rating gap - prioritize user satisfaction improvements") + actions.append("Analyze negative reviews to identify top issues") + actions.append("Implement in-app rating prompts after positive experiences") + actions.append("Respond to all negative reviews professionally") + elif rating_gap > 0.2: + actions.append("Focus on incremental improvements to close rating gap") + actions.append("Optimize timing of rating requests") + else: + actions.append("Ratings are competitive - maintain quality and continue improvements") + + return actions + + def _generate_content_recommendations(self, desc_length_gap: int) -> List[str]: + """Generate content recommendations based on length gap.""" + recommendations = [] + + if desc_length_gap > 500: + recommendations.append( + "Expand description to match competitor detail level" + ) + recommendations.append( + "Add use case examples and success stories" + ) + recommendations.append( + "Include more feature explanations and benefits" + ) + elif desc_length_gap < -500: + recommendations.append( + "Consider condensing description for better readability" + ) + recommendations.append( + "Focus on most important features first" + ) + else: + recommendations.append( + "Description length is competitive" + ) + + return recommendations + + def _assess_competitive_position( + self, + your_analysis: Dict[str, Any], + competitor_comparison: Dict[str, Any] + ) -> str: + """Assess your competitive position.""" + your_strength = your_analysis['competitive_strength'] + competitors = competitor_comparison['ranked_competitors'] + + if not competitors: + return "No comparison data available" + + # Find where you'd rank + better_than_count = sum( + 1 for comp in competitors + if your_strength > comp['competitive_strength'] + ) + + position_percentage = (better_than_count / len(competitors)) * 100 + + if position_percentage >= 75: + return "Strong Position: Top quartile in competitive strength" + elif position_percentage >= 50: + return "Competitive Position: Above average, opportunities for improvement" + elif position_percentage >= 25: + return "Challenging Position: Below average, requires strategic improvements" + else: + return "Weak Position: Bottom quartile, major ASO overhaul needed" + + +def analyze_competitor_set( + category: str, + competitors_data: List[Dict[str, Any]], + platform: str = 'apple' +) -> Dict[str, Any]: + """ + Convenience function to analyze a set of competitors. + + Args: + category: App category + competitors_data: List of competitor data + platform: 'apple' or 'google' + + Returns: + Complete competitive analysis + """ + analyzer = CompetitorAnalyzer(category, platform) + return analyzer.compare_competitors(competitors_data) diff --git a/skills/app-store-optimization/scripts/keyword_analyzer.py b/skills/app-store-optimization/scripts/keyword_analyzer.py new file mode 100644 index 00000000..5c3d80ba --- /dev/null +++ b/skills/app-store-optimization/scripts/keyword_analyzer.py @@ -0,0 +1,406 @@ +""" +Keyword analysis module for App Store Optimization. +Analyzes keyword search volume, competition, and relevance for app discovery. +""" + +from typing import Dict, List, Any, Optional, Tuple +import re +from collections import Counter + + +class KeywordAnalyzer: + """Analyzes keywords for ASO effectiveness.""" + + # Competition level thresholds (based on number of competing apps) + COMPETITION_THRESHOLDS = { + 'low': 1000, + 'medium': 5000, + 'high': 10000 + } + + # Search volume categories (monthly searches estimate) + VOLUME_CATEGORIES = { + 'very_low': 1000, + 'low': 5000, + 'medium': 20000, + 'high': 100000, + 'very_high': 500000 + } + + def __init__(self): + """Initialize keyword analyzer.""" + self.analyzed_keywords = {} + + def analyze_keyword( + self, + keyword: str, + search_volume: int = 0, + competing_apps: int = 0, + relevance_score: float = 0.0 + ) -> Dict[str, Any]: + """ + Analyze a single keyword for ASO potential. + + Args: + keyword: The keyword to analyze + search_volume: Estimated monthly search volume + competing_apps: Number of apps competing for this keyword + relevance_score: Relevance to your app (0.0-1.0) + + Returns: + Dictionary with keyword analysis + """ + competition_level = self._calculate_competition_level(competing_apps) + volume_category = self._categorize_search_volume(search_volume) + difficulty_score = self._calculate_keyword_difficulty( + search_volume, + competing_apps + ) + + # Calculate potential score (0-100) + potential_score = self._calculate_potential_score( + search_volume, + competing_apps, + relevance_score + ) + + analysis = { + 'keyword': keyword, + 'search_volume': search_volume, + 'volume_category': volume_category, + 'competing_apps': competing_apps, + 'competition_level': competition_level, + 'relevance_score': relevance_score, + 'difficulty_score': difficulty_score, + 'potential_score': potential_score, + 'recommendation': self._generate_recommendation( + potential_score, + difficulty_score, + relevance_score + ), + 'keyword_length': len(keyword.split()), + 'is_long_tail': len(keyword.split()) >= 3 + } + + self.analyzed_keywords[keyword] = analysis + return analysis + + def compare_keywords(self, keywords_data: List[Dict[str, Any]]) -> Dict[str, Any]: + """ + Compare multiple keywords and rank by potential. + + Args: + keywords_data: List of dicts with keyword, search_volume, competing_apps, relevance_score + + Returns: + Comparison report with ranked keywords + """ + analyses = [] + for kw_data in keywords_data: + analysis = self.analyze_keyword( + keyword=kw_data['keyword'], + search_volume=kw_data.get('search_volume', 0), + competing_apps=kw_data.get('competing_apps', 0), + relevance_score=kw_data.get('relevance_score', 0.0) + ) + analyses.append(analysis) + + # Sort by potential score (descending) + ranked_keywords = sorted( + analyses, + key=lambda x: x['potential_score'], + reverse=True + ) + + # Categorize keywords + primary_keywords = [ + kw for kw in ranked_keywords + if kw['potential_score'] >= 70 and kw['relevance_score'] >= 0.8 + ] + + secondary_keywords = [ + kw for kw in ranked_keywords + if 50 <= kw['potential_score'] < 70 and kw['relevance_score'] >= 0.6 + ] + + long_tail_keywords = [ + kw for kw in ranked_keywords + if kw['is_long_tail'] and kw['relevance_score'] >= 0.7 + ] + + return { + 'total_keywords_analyzed': len(analyses), + 'ranked_keywords': ranked_keywords, + 'primary_keywords': primary_keywords[:5], # Top 5 + 'secondary_keywords': secondary_keywords[:10], # Top 10 + 'long_tail_keywords': long_tail_keywords[:10], # Top 10 + 'summary': self._generate_comparison_summary( + primary_keywords, + secondary_keywords, + long_tail_keywords + ) + } + + def find_long_tail_opportunities( + self, + base_keyword: str, + modifiers: List[str] + ) -> List[Dict[str, Any]]: + """ + Generate long-tail keyword variations. + + Args: + base_keyword: Core keyword (e.g., "task manager") + modifiers: List of modifiers (e.g., ["free", "simple", "team"]) + + Returns: + List of long-tail keyword suggestions + """ + long_tail_keywords = [] + + # Generate combinations + for modifier in modifiers: + # Modifier + base + variation1 = f"{modifier} {base_keyword}" + long_tail_keywords.append({ + 'keyword': variation1, + 'pattern': 'modifier_base', + 'estimated_competition': 'low', + 'rationale': f"Less competitive variation of '{base_keyword}'" + }) + + # Base + modifier + variation2 = f"{base_keyword} {modifier}" + long_tail_keywords.append({ + 'keyword': variation2, + 'pattern': 'base_modifier', + 'estimated_competition': 'low', + 'rationale': f"Specific use-case variation of '{base_keyword}'" + }) + + # Add question-based long-tail + question_words = ['how', 'what', 'best', 'top'] + for q_word in question_words: + question_keyword = f"{q_word} {base_keyword}" + long_tail_keywords.append({ + 'keyword': question_keyword, + 'pattern': 'question_based', + 'estimated_competition': 'very_low', + 'rationale': f"Informational search query" + }) + + return long_tail_keywords + + def extract_keywords_from_text( + self, + text: str, + min_word_length: int = 3 + ) -> List[Tuple[str, int]]: + """ + Extract potential keywords from text (descriptions, reviews). + + Args: + text: Text to analyze + min_word_length: Minimum word length to consider + + Returns: + List of (keyword, frequency) tuples + """ + # Clean and normalize text + text = text.lower() + text = re.sub(r'[^\w\s]', ' ', text) + + # Extract words + words = text.split() + + # Filter by length + words = [w for w in words if len(w) >= min_word_length] + + # Remove common stop words + stop_words = { + 'the', 'and', 'for', 'with', 'this', 'that', 'from', 'have', + 'but', 'not', 'you', 'all', 'can', 'are', 'was', 'were', 'been' + } + words = [w for w in words if w not in stop_words] + + # Count frequency + word_counts = Counter(words) + + # Extract 2-word phrases + phrases = [] + for i in range(len(words) - 1): + phrase = f"{words[i]} {words[i+1]}" + phrases.append(phrase) + + phrase_counts = Counter(phrases) + + # Combine and sort + all_keywords = list(word_counts.items()) + list(phrase_counts.items()) + all_keywords.sort(key=lambda x: x[1], reverse=True) + + return all_keywords[:50] # Top 50 + + def calculate_keyword_density( + self, + text: str, + target_keywords: List[str] + ) -> Dict[str, float]: + """ + Calculate keyword density in text. + + Args: + text: Text to analyze (title, description) + target_keywords: Keywords to check density for + + Returns: + Dictionary of keyword: density (percentage) + """ + text_lower = text.lower() + total_words = len(text_lower.split()) + + densities = {} + for keyword in target_keywords: + keyword_lower = keyword.lower() + occurrences = text_lower.count(keyword_lower) + density = (occurrences / total_words) * 100 if total_words > 0 else 0 + densities[keyword] = round(density, 2) + + return densities + + def _calculate_competition_level(self, competing_apps: int) -> str: + """Determine competition level based on number of competing apps.""" + if competing_apps < self.COMPETITION_THRESHOLDS['low']: + return 'low' + elif competing_apps < self.COMPETITION_THRESHOLDS['medium']: + return 'medium' + elif competing_apps < self.COMPETITION_THRESHOLDS['high']: + return 'high' + else: + return 'very_high' + + def _categorize_search_volume(self, search_volume: int) -> str: + """Categorize search volume.""" + if search_volume < self.VOLUME_CATEGORIES['very_low']: + return 'very_low' + elif search_volume < self.VOLUME_CATEGORIES['low']: + return 'low' + elif search_volume < self.VOLUME_CATEGORIES['medium']: + return 'medium' + elif search_volume < self.VOLUME_CATEGORIES['high']: + return 'high' + else: + return 'very_high' + + def _calculate_keyword_difficulty( + self, + search_volume: int, + competing_apps: int + ) -> float: + """ + Calculate keyword difficulty score (0-100). + Higher score = harder to rank. + """ + if competing_apps == 0: + return 0.0 + + # Competition factor (0-1) + competition_factor = min(competing_apps / 50000, 1.0) + + # Volume factor (0-1) - higher volume = more difficulty + volume_factor = min(search_volume / 1000000, 1.0) + + # Difficulty score (weighted average) + difficulty = (competition_factor * 0.7 + volume_factor * 0.3) * 100 + + return round(difficulty, 1) + + def _calculate_potential_score( + self, + search_volume: int, + competing_apps: int, + relevance_score: float + ) -> float: + """ + Calculate overall keyword potential (0-100). + Higher score = better opportunity. + """ + # Volume score (0-40 points) + volume_score = min((search_volume / 100000) * 40, 40) + + # Competition score (0-30 points) - inverse relationship + if competing_apps > 0: + competition_score = max(30 - (competing_apps / 500), 0) + else: + competition_score = 30 + + # Relevance score (0-30 points) + relevance_points = relevance_score * 30 + + total_score = volume_score + competition_score + relevance_points + + return round(min(total_score, 100), 1) + + def _generate_recommendation( + self, + potential_score: float, + difficulty_score: float, + relevance_score: float + ) -> str: + """Generate actionable recommendation for keyword.""" + if relevance_score < 0.5: + return "Low relevance - avoid targeting" + + if potential_score >= 70: + return "High priority - target immediately" + elif potential_score >= 50: + if difficulty_score < 50: + return "Good opportunity - include in metadata" + else: + return "Competitive - use in description, not title" + elif potential_score >= 30: + return "Secondary keyword - use for long-tail variations" + else: + return "Low potential - deprioritize" + + def _generate_comparison_summary( + self, + primary_keywords: List[Dict[str, Any]], + secondary_keywords: List[Dict[str, Any]], + long_tail_keywords: List[Dict[str, Any]] + ) -> str: + """Generate summary of keyword comparison.""" + summary_parts = [] + + summary_parts.append( + f"Identified {len(primary_keywords)} high-priority primary keywords." + ) + + if primary_keywords: + top_keyword = primary_keywords[0]['keyword'] + summary_parts.append( + f"Top recommendation: '{top_keyword}' (potential score: {primary_keywords[0]['potential_score']})." + ) + + summary_parts.append( + f"Found {len(secondary_keywords)} secondary keywords for description and metadata." + ) + + summary_parts.append( + f"Discovered {len(long_tail_keywords)} long-tail opportunities with lower competition." + ) + + return " ".join(summary_parts) + + +def analyze_keyword_set(keywords_data: List[Dict[str, Any]]) -> Dict[str, Any]: + """ + Convenience function to analyze a set of keywords. + + Args: + keywords_data: List of keyword data dictionaries + + Returns: + Complete analysis report + """ + analyzer = KeywordAnalyzer() + return analyzer.compare_keywords(keywords_data) diff --git a/skills/app-store-optimization/scripts/launch_checklist.py b/skills/app-store-optimization/scripts/launch_checklist.py new file mode 100644 index 00000000..38eea18b --- /dev/null +++ b/skills/app-store-optimization/scripts/launch_checklist.py @@ -0,0 +1,739 @@ +""" +Launch checklist module for App Store Optimization. +Generates comprehensive pre-launch and update checklists. +""" + +from typing import Dict, List, Any, Optional +from datetime import datetime, timedelta + + +class LaunchChecklistGenerator: + """Generates comprehensive checklists for app launches and updates.""" + + def __init__(self, platform: str = 'both'): + """ + Initialize checklist generator. + + Args: + platform: 'apple', 'google', or 'both' + """ + if platform not in ['apple', 'google', 'both']: + raise ValueError("Platform must be 'apple', 'google', or 'both'") + + self.platform = platform + + def generate_prelaunch_checklist( + self, + app_info: Dict[str, Any], + launch_date: Optional[str] = None + ) -> Dict[str, Any]: + """ + Generate comprehensive pre-launch checklist. + + Args: + app_info: App information (name, category, target_audience) + launch_date: Target launch date (YYYY-MM-DD) + + Returns: + Complete pre-launch checklist + """ + checklist = { + 'app_info': app_info, + 'launch_date': launch_date, + 'checklists': {} + } + + # Generate platform-specific checklists + if self.platform in ['apple', 'both']: + checklist['checklists']['apple'] = self._generate_apple_checklist(app_info) + + if self.platform in ['google', 'both']: + checklist['checklists']['google'] = self._generate_google_checklist(app_info) + + # Add universal checklist items + checklist['checklists']['universal'] = self._generate_universal_checklist(app_info) + + # Generate timeline + if launch_date: + checklist['timeline'] = self._generate_launch_timeline(launch_date) + + # Calculate completion status + checklist['summary'] = self._calculate_checklist_summary(checklist['checklists']) + + return checklist + + def validate_app_store_compliance( + self, + app_data: Dict[str, Any], + platform: str = 'apple' + ) -> Dict[str, Any]: + """ + Validate compliance with app store guidelines. + + Args: + app_data: App data including metadata, privacy policy, etc. + platform: 'apple' or 'google' + + Returns: + Compliance validation report + """ + validation_results = { + 'platform': platform, + 'is_compliant': True, + 'errors': [], + 'warnings': [], + 'recommendations': [] + } + + if platform == 'apple': + self._validate_apple_compliance(app_data, validation_results) + elif platform == 'google': + self._validate_google_compliance(app_data, validation_results) + + # Determine overall compliance + validation_results['is_compliant'] = len(validation_results['errors']) == 0 + + return validation_results + + def create_update_plan( + self, + current_version: str, + planned_features: List[str], + update_frequency: str = 'monthly' + ) -> Dict[str, Any]: + """ + Create update cadence and feature rollout plan. + + Args: + current_version: Current app version + planned_features: List of planned features + update_frequency: 'weekly', 'biweekly', 'monthly', 'quarterly' + + Returns: + Update plan with cadence and feature schedule + """ + # Calculate next versions + next_versions = self._calculate_next_versions( + current_version, + update_frequency, + len(planned_features) + ) + + # Distribute features across versions + feature_schedule = self._distribute_features( + planned_features, + next_versions + ) + + # Generate "What's New" templates + whats_new_templates = [ + self._generate_whats_new_template(version_data) + for version_data in feature_schedule + ] + + return { + 'current_version': current_version, + 'update_frequency': update_frequency, + 'planned_updates': len(feature_schedule), + 'feature_schedule': feature_schedule, + 'whats_new_templates': whats_new_templates, + 'recommendations': self._generate_update_recommendations(update_frequency) + } + + def optimize_launch_timing( + self, + app_category: str, + target_audience: str, + current_date: Optional[str] = None + ) -> Dict[str, Any]: + """ + Recommend optimal launch timing. + + Args: + app_category: App category + target_audience: Target audience description + current_date: Current date (YYYY-MM-DD), defaults to today + + Returns: + Launch timing recommendations + """ + if not current_date: + current_date = datetime.now().strftime('%Y-%m-%d') + + # Analyze launch timing factors + day_of_week_rec = self._recommend_day_of_week(app_category) + seasonal_rec = self._recommend_seasonal_timing(app_category, current_date) + competitive_rec = self._analyze_competitive_timing(app_category) + + # Calculate optimal dates + optimal_dates = self._calculate_optimal_dates( + current_date, + day_of_week_rec, + seasonal_rec + ) + + return { + 'current_date': current_date, + 'optimal_launch_dates': optimal_dates, + 'day_of_week_recommendation': day_of_week_rec, + 'seasonal_considerations': seasonal_rec, + 'competitive_timing': competitive_rec, + 'final_recommendation': self._generate_timing_recommendation( + optimal_dates, + seasonal_rec + ) + } + + def plan_seasonal_campaigns( + self, + app_category: str, + current_month: int = None + ) -> Dict[str, Any]: + """ + Identify seasonal opportunities for ASO campaigns. + + Args: + app_category: App category + current_month: Current month (1-12), defaults to current + + Returns: + Seasonal campaign opportunities + """ + if not current_month: + current_month = datetime.now().month + + # Identify relevant seasonal events + seasonal_opportunities = self._identify_seasonal_opportunities( + app_category, + current_month + ) + + # Generate campaign ideas + campaigns = [ + self._generate_seasonal_campaign(opportunity) + for opportunity in seasonal_opportunities + ] + + return { + 'current_month': current_month, + 'category': app_category, + 'seasonal_opportunities': seasonal_opportunities, + 'campaign_ideas': campaigns, + 'implementation_timeline': self._create_seasonal_timeline(campaigns) + } + + def _generate_apple_checklist(self, app_info: Dict[str, Any]) -> List[Dict[str, Any]]: + """Generate Apple App Store specific checklist.""" + return [ + { + 'category': 'App Store Connect Setup', + 'items': [ + {'task': 'App Store Connect account created', 'status': 'pending'}, + {'task': 'App bundle ID registered', 'status': 'pending'}, + {'task': 'App Privacy declarations completed', 'status': 'pending'}, + {'task': 'Age rating questionnaire completed', 'status': 'pending'} + ] + }, + { + 'category': 'Metadata (Apple)', + 'items': [ + {'task': 'App title (30 chars max)', 'status': 'pending'}, + {'task': 'Subtitle (30 chars max)', 'status': 'pending'}, + {'task': 'Promotional text (170 chars max)', 'status': 'pending'}, + {'task': 'Description (4000 chars max)', 'status': 'pending'}, + {'task': 'Keywords (100 chars, comma-separated)', 'status': 'pending'}, + {'task': 'Category selection (primary + secondary)', 'status': 'pending'} + ] + }, + { + 'category': 'Visual Assets (Apple)', + 'items': [ + {'task': 'App icon (1024x1024px)', 'status': 'pending'}, + {'task': 'Screenshots (iPhone 6.7" required)', 'status': 'pending'}, + {'task': 'Screenshots (iPhone 5.5" required)', 'status': 'pending'}, + {'task': 'Screenshots (iPad Pro 12.9" if iPad app)', 'status': 'pending'}, + {'task': 'App preview video (optional but recommended)', 'status': 'pending'} + ] + }, + { + 'category': 'Technical Requirements (Apple)', + 'items': [ + {'task': 'Build uploaded to App Store Connect', 'status': 'pending'}, + {'task': 'TestFlight testing completed', 'status': 'pending'}, + {'task': 'App tested on required iOS versions', 'status': 'pending'}, + {'task': 'Crash-free rate > 99%', 'status': 'pending'}, + {'task': 'All links in app/metadata working', 'status': 'pending'} + ] + }, + { + 'category': 'Legal & Privacy (Apple)', + 'items': [ + {'task': 'Privacy Policy URL provided', 'status': 'pending'}, + {'task': 'Terms of Service URL (if applicable)', 'status': 'pending'}, + {'task': 'Data collection declarations accurate', 'status': 'pending'}, + {'task': 'Third-party SDKs disclosed', 'status': 'pending'} + ] + } + ] + + def _generate_google_checklist(self, app_info: Dict[str, Any]) -> List[Dict[str, Any]]: + """Generate Google Play Store specific checklist.""" + return [ + { + 'category': 'Play Console Setup', + 'items': [ + {'task': 'Google Play Console account created', 'status': 'pending'}, + {'task': 'Developer profile completed', 'status': 'pending'}, + {'task': 'Payment merchant account linked (if paid app)', 'status': 'pending'}, + {'task': 'Content rating questionnaire completed', 'status': 'pending'} + ] + }, + { + 'category': 'Metadata (Google)', + 'items': [ + {'task': 'App title (50 chars max)', 'status': 'pending'}, + {'task': 'Short description (80 chars max)', 'status': 'pending'}, + {'task': 'Full description (4000 chars max)', 'status': 'pending'}, + {'task': 'Category selection', 'status': 'pending'}, + {'task': 'Tags (up to 5)', 'status': 'pending'} + ] + }, + { + 'category': 'Visual Assets (Google)', + 'items': [ + {'task': 'App icon (512x512px)', 'status': 'pending'}, + {'task': 'Feature graphic (1024x500px)', 'status': 'pending'}, + {'task': 'Screenshots (2-8 required, phone)', 'status': 'pending'}, + {'task': 'Screenshots (tablet, if applicable)', 'status': 'pending'}, + {'task': 'Promo video (YouTube link, optional)', 'status': 'pending'} + ] + }, + { + 'category': 'Technical Requirements (Google)', + 'items': [ + {'task': 'APK/AAB uploaded to Play Console', 'status': 'pending'}, + {'task': 'Internal testing completed', 'status': 'pending'}, + {'task': 'App tested on required Android versions', 'status': 'pending'}, + {'task': 'Target API level meets requirements', 'status': 'pending'}, + {'task': 'All permissions justified', 'status': 'pending'} + ] + }, + { + 'category': 'Legal & Privacy (Google)', + 'items': [ + {'task': 'Privacy Policy URL provided', 'status': 'pending'}, + {'task': 'Data safety section completed', 'status': 'pending'}, + {'task': 'Ads disclosure (if applicable)', 'status': 'pending'}, + {'task': 'In-app purchase disclosure (if applicable)', 'status': 'pending'} + ] + } + ] + + def _generate_universal_checklist(self, app_info: Dict[str, Any]) -> List[Dict[str, Any]]: + """Generate universal (both platforms) checklist.""" + return [ + { + 'category': 'Pre-Launch Marketing', + 'items': [ + {'task': 'Landing page created', 'status': 'pending'}, + {'task': 'Social media accounts setup', 'status': 'pending'}, + {'task': 'Press kit prepared', 'status': 'pending'}, + {'task': 'Beta tester feedback collected', 'status': 'pending'}, + {'task': 'Launch announcement drafted', 'status': 'pending'} + ] + }, + { + 'category': 'ASO Preparation', + 'items': [ + {'task': 'Keyword research completed', 'status': 'pending'}, + {'task': 'Competitor analysis done', 'status': 'pending'}, + {'task': 'A/B test plan created for post-launch', 'status': 'pending'}, + {'task': 'Analytics tracking configured', 'status': 'pending'} + ] + }, + { + 'category': 'Quality Assurance', + 'items': [ + {'task': 'All core features tested', 'status': 'pending'}, + {'task': 'User flows validated', 'status': 'pending'}, + {'task': 'Performance testing completed', 'status': 'pending'}, + {'task': 'Accessibility features tested', 'status': 'pending'}, + {'task': 'Security audit completed', 'status': 'pending'} + ] + }, + { + 'category': 'Support Infrastructure', + 'items': [ + {'task': 'Support email/system setup', 'status': 'pending'}, + {'task': 'FAQ page created', 'status': 'pending'}, + {'task': 'Documentation for users prepared', 'status': 'pending'}, + {'task': 'Team trained on handling reviews', 'status': 'pending'} + ] + } + ] + + def _generate_launch_timeline(self, launch_date: str) -> List[Dict[str, Any]]: + """Generate timeline with milestones leading to launch.""" + launch_dt = datetime.strptime(launch_date, '%Y-%m-%d') + + milestones = [ + { + 'date': (launch_dt - timedelta(days=90)).strftime('%Y-%m-%d'), + 'milestone': '90 days before: Complete keyword research and competitor analysis' + }, + { + 'date': (launch_dt - timedelta(days=60)).strftime('%Y-%m-%d'), + 'milestone': '60 days before: Finalize metadata and visual assets' + }, + { + 'date': (launch_dt - timedelta(days=45)).strftime('%Y-%m-%d'), + 'milestone': '45 days before: Begin beta testing program' + }, + { + 'date': (launch_dt - timedelta(days=30)).strftime('%Y-%m-%d'), + 'milestone': '30 days before: Submit app for review (Apple typically takes 1-2 days, Google instant)' + }, + { + 'date': (launch_dt - timedelta(days=14)).strftime('%Y-%m-%d'), + 'milestone': '14 days before: Prepare launch marketing materials' + }, + { + 'date': (launch_dt - timedelta(days=7)).strftime('%Y-%m-%d'), + 'milestone': '7 days before: Set up analytics and monitoring' + }, + { + 'date': launch_dt.strftime('%Y-%m-%d'), + 'milestone': 'Launch Day: Release app and execute marketing plan' + }, + { + 'date': (launch_dt + timedelta(days=7)).strftime('%Y-%m-%d'), + 'milestone': '7 days after: Monitor metrics, respond to reviews, address critical issues' + }, + { + 'date': (launch_dt + timedelta(days=30)).strftime('%Y-%m-%d'), + 'milestone': '30 days after: Analyze launch metrics, plan first update' + } + ] + + return milestones + + def _calculate_checklist_summary(self, checklists: Dict[str, List[Dict[str, Any]]]) -> Dict[str, Any]: + """Calculate completion summary.""" + total_items = 0 + completed_items = 0 + + for platform, categories in checklists.items(): + for category in categories: + for item in category['items']: + total_items += 1 + if item['status'] == 'completed': + completed_items += 1 + + completion_percentage = (completed_items / total_items * 100) if total_items > 0 else 0 + + return { + 'total_items': total_items, + 'completed_items': completed_items, + 'pending_items': total_items - completed_items, + 'completion_percentage': round(completion_percentage, 1), + 'is_ready_to_launch': completion_percentage == 100 + } + + def _validate_apple_compliance( + self, + app_data: Dict[str, Any], + validation_results: Dict[str, Any] + ) -> None: + """Validate Apple App Store compliance.""" + # Check for required fields + if not app_data.get('privacy_policy_url'): + validation_results['errors'].append("Privacy Policy URL is required") + + if not app_data.get('app_icon'): + validation_results['errors'].append("App icon (1024x1024px) is required") + + # Check metadata character limits + title = app_data.get('title', '') + if len(title) > 30: + validation_results['errors'].append(f"Title exceeds 30 characters ({len(title)})") + + # Warnings for best practices + subtitle = app_data.get('subtitle', '') + if not subtitle: + validation_results['warnings'].append("Subtitle is empty - consider adding for better discoverability") + + keywords = app_data.get('keywords', '') + if len(keywords) < 80: + validation_results['warnings'].append( + f"Keywords field underutilized ({len(keywords)}/100 chars) - add more keywords" + ) + + def _validate_google_compliance( + self, + app_data: Dict[str, Any], + validation_results: Dict[str, Any] + ) -> None: + """Validate Google Play Store compliance.""" + # Check for required fields + if not app_data.get('privacy_policy_url'): + validation_results['errors'].append("Privacy Policy URL is required") + + if not app_data.get('feature_graphic'): + validation_results['errors'].append("Feature graphic (1024x500px) is required") + + # Check metadata character limits + title = app_data.get('title', '') + if len(title) > 50: + validation_results['errors'].append(f"Title exceeds 50 characters ({len(title)})") + + short_desc = app_data.get('short_description', '') + if len(short_desc) > 80: + validation_results['errors'].append(f"Short description exceeds 80 characters ({len(short_desc)})") + + # Warnings + if not short_desc: + validation_results['warnings'].append("Short description is empty") + + def _calculate_next_versions( + self, + current_version: str, + update_frequency: str, + feature_count: int + ) -> List[str]: + """Calculate next version numbers.""" + # Parse current version (assume semantic versioning) + parts = current_version.split('.') + major, minor, patch = int(parts[0]), int(parts[1]), int(parts[2] if len(parts) > 2 else 0) + + versions = [] + for i in range(feature_count): + if update_frequency == 'weekly': + patch += 1 + elif update_frequency == 'biweekly': + patch += 1 + elif update_frequency == 'monthly': + minor += 1 + patch = 0 + else: # quarterly + minor += 1 + patch = 0 + + versions.append(f"{major}.{minor}.{patch}") + + return versions + + def _distribute_features( + self, + features: List[str], + versions: List[str] + ) -> List[Dict[str, Any]]: + """Distribute features across versions.""" + features_per_version = max(1, len(features) // len(versions)) + + schedule = [] + for i, version in enumerate(versions): + start_idx = i * features_per_version + end_idx = start_idx + features_per_version if i < len(versions) - 1 else len(features) + + schedule.append({ + 'version': version, + 'features': features[start_idx:end_idx], + 'release_priority': 'high' if i == 0 else ('medium' if i < len(versions) // 2 else 'low') + }) + + return schedule + + def _generate_whats_new_template(self, version_data: Dict[str, Any]) -> Dict[str, str]: + """Generate What's New template for version.""" + features_list = '\n'.join([f"• {feature}" for feature in version_data['features']]) + + template = f"""Version {version_data['version']} + +{features_list} + +We're constantly improving your experience. Thanks for using [App Name]! + +Have feedback? Contact us at support@[company].com""" + + return { + 'version': version_data['version'], + 'template': template + } + + def _generate_update_recommendations(self, update_frequency: str) -> List[str]: + """Generate recommendations for update strategy.""" + recommendations = [] + + if update_frequency == 'weekly': + recommendations.append("Weekly updates show active development but ensure quality doesn't suffer") + elif update_frequency == 'monthly': + recommendations.append("Monthly updates are optimal for most apps - balance features and stability") + + recommendations.extend([ + "Include bug fixes in every update", + "Update 'What's New' section with each release", + "Respond to reviews mentioning fixed issues" + ]) + + return recommendations + + def _recommend_day_of_week(self, app_category: str) -> Dict[str, Any]: + """Recommend best day of week to launch.""" + # General recommendations based on category + if app_category.lower() in ['games', 'entertainment']: + return { + 'recommended_day': 'Thursday', + 'rationale': 'People download entertainment apps before weekend' + } + elif app_category.lower() in ['productivity', 'business']: + return { + 'recommended_day': 'Tuesday', + 'rationale': 'Business users most active mid-week' + } + else: + return { + 'recommended_day': 'Wednesday', + 'rationale': 'Mid-week provides good balance and review potential' + } + + def _recommend_seasonal_timing(self, app_category: str, current_date: str) -> Dict[str, Any]: + """Recommend seasonal timing considerations.""" + current_dt = datetime.strptime(current_date, '%Y-%m-%d') + month = current_dt.month + + # Avoid certain periods + avoid_periods = [] + if month == 12: + avoid_periods.append("Late December - low user engagement during holidays") + if month in [7, 8]: + avoid_periods.append("Summer months - some categories see lower engagement") + + # Recommend periods + good_periods = [] + if month in [1, 9]: + good_periods.append("New Year/Back-to-school - high user engagement") + if month in [10, 11]: + good_periods.append("Pre-holiday season - good for shopping/gift apps") + + return { + 'current_month': month, + 'avoid_periods': avoid_periods, + 'good_periods': good_periods + } + + def _analyze_competitive_timing(self, app_category: str) -> Dict[str, str]: + """Analyze competitive timing considerations.""" + return { + 'recommendation': 'Research competitor launch schedules in your category', + 'strategy': 'Avoid launching same week as major competitor updates' + } + + def _calculate_optimal_dates( + self, + current_date: str, + day_rec: Dict[str, Any], + seasonal_rec: Dict[str, Any] + ) -> List[str]: + """Calculate optimal launch dates.""" + current_dt = datetime.strptime(current_date, '%Y-%m-%d') + + # Find next occurrence of recommended day + target_day = day_rec['recommended_day'] + days_map = {'Monday': 0, 'Tuesday': 1, 'Wednesday': 2, 'Thursday': 3, 'Friday': 4} + target_day_num = days_map.get(target_day, 2) + + days_ahead = (target_day_num - current_dt.weekday()) % 7 + if days_ahead == 0: + days_ahead = 7 + + next_target_date = current_dt + timedelta(days=days_ahead) + + optimal_dates = [ + next_target_date.strftime('%Y-%m-%d'), + (next_target_date + timedelta(days=7)).strftime('%Y-%m-%d'), + (next_target_date + timedelta(days=14)).strftime('%Y-%m-%d') + ] + + return optimal_dates + + def _generate_timing_recommendation( + self, + optimal_dates: List[str], + seasonal_rec: Dict[str, Any] + ) -> str: + """Generate final timing recommendation.""" + if seasonal_rec['avoid_periods']: + return f"Consider launching in {optimal_dates[1]} to avoid {seasonal_rec['avoid_periods'][0]}" + elif seasonal_rec['good_periods']: + return f"Launch on {optimal_dates[0]} to capitalize on {seasonal_rec['good_periods'][0]}" + else: + return f"Recommended launch date: {optimal_dates[0]}" + + def _identify_seasonal_opportunities( + self, + app_category: str, + current_month: int + ) -> List[Dict[str, Any]]: + """Identify seasonal opportunities for category.""" + opportunities = [] + + # Universal opportunities + if current_month == 1: + opportunities.append({ + 'event': 'New Year Resolutions', + 'dates': 'January 1-31', + 'relevance': 'high' if app_category.lower() in ['health', 'fitness', 'productivity'] else 'medium' + }) + + if current_month in [11, 12]: + opportunities.append({ + 'event': 'Holiday Shopping Season', + 'dates': 'November-December', + 'relevance': 'high' if app_category.lower() in ['shopping', 'gifts'] else 'low' + }) + + # Category-specific + if app_category.lower() == 'education' and current_month in [8, 9]: + opportunities.append({ + 'event': 'Back to School', + 'dates': 'August-September', + 'relevance': 'high' + }) + + return opportunities + + def _generate_seasonal_campaign(self, opportunity: Dict[str, Any]) -> Dict[str, Any]: + """Generate campaign idea for seasonal opportunity.""" + return { + 'event': opportunity['event'], + 'campaign_idea': f"Create themed visuals and messaging for {opportunity['event']}", + 'metadata_updates': 'Update app description and screenshots with seasonal themes', + 'promotion_strategy': 'Consider limited-time features or discounts' + } + + def _create_seasonal_timeline(self, campaigns: List[Dict[str, Any]]) -> List[str]: + """Create implementation timeline for campaigns.""" + return [ + f"30 days before: Plan {campaign['event']} campaign strategy" + for campaign in campaigns + ] + + +def generate_launch_checklist( + platform: str, + app_info: Dict[str, Any], + launch_date: Optional[str] = None +) -> Dict[str, Any]: + """ + Convenience function to generate launch checklist. + + Args: + platform: Platform ('apple', 'google', or 'both') + app_info: App information + launch_date: Target launch date + + Returns: + Complete launch checklist + """ + generator = LaunchChecklistGenerator(platform) + return generator.generate_prelaunch_checklist(app_info, launch_date) diff --git a/skills/app-store-optimization/scripts/localization_helper.py b/skills/app-store-optimization/scripts/localization_helper.py new file mode 100644 index 00000000..c47003ca --- /dev/null +++ b/skills/app-store-optimization/scripts/localization_helper.py @@ -0,0 +1,588 @@ +""" +Localization helper module for App Store Optimization. +Manages multi-language ASO optimization strategies. +""" + +from typing import Dict, List, Any, Optional, Tuple + + +class LocalizationHelper: + """Helps manage multi-language ASO optimization.""" + + # Priority markets by language (based on app store revenue and user base) + PRIORITY_MARKETS = { + 'tier_1': [ + {'language': 'en-US', 'market': 'United States', 'revenue_share': 0.25}, + {'language': 'zh-CN', 'market': 'China', 'revenue_share': 0.20}, + {'language': 'ja-JP', 'market': 'Japan', 'revenue_share': 0.10}, + {'language': 'de-DE', 'market': 'Germany', 'revenue_share': 0.08}, + {'language': 'en-GB', 'market': 'United Kingdom', 'revenue_share': 0.06} + ], + 'tier_2': [ + {'language': 'fr-FR', 'market': 'France', 'revenue_share': 0.05}, + {'language': 'ko-KR', 'market': 'South Korea', 'revenue_share': 0.05}, + {'language': 'es-ES', 'market': 'Spain', 'revenue_share': 0.03}, + {'language': 'it-IT', 'market': 'Italy', 'revenue_share': 0.03}, + {'language': 'pt-BR', 'market': 'Brazil', 'revenue_share': 0.03} + ], + 'tier_3': [ + {'language': 'ru-RU', 'market': 'Russia', 'revenue_share': 0.02}, + {'language': 'es-MX', 'market': 'Mexico', 'revenue_share': 0.02}, + {'language': 'nl-NL', 'market': 'Netherlands', 'revenue_share': 0.02}, + {'language': 'sv-SE', 'market': 'Sweden', 'revenue_share': 0.01}, + {'language': 'pl-PL', 'market': 'Poland', 'revenue_share': 0.01} + ] + } + + # Character limit multipliers by language (some languages need more/less space) + CHAR_MULTIPLIERS = { + 'en': 1.0, + 'zh': 0.6, # Chinese characters are more compact + 'ja': 0.7, # Japanese uses kanji + 'ko': 0.8, # Korean is relatively compact + 'de': 1.3, # German words are typically longer + 'fr': 1.2, # French tends to be longer + 'es': 1.1, # Spanish slightly longer + 'pt': 1.1, # Portuguese similar to Spanish + 'ru': 1.1, # Russian similar length + 'ar': 1.0, # Arabic varies + 'it': 1.1 # Italian similar to Spanish + } + + def __init__(self, app_category: str = 'general'): + """ + Initialize localization helper. + + Args: + app_category: App category to prioritize relevant markets + """ + self.app_category = app_category + self.localization_plans = [] + + def identify_target_markets( + self, + current_market: str = 'en-US', + budget_level: str = 'medium', + target_market_count: int = 5 + ) -> Dict[str, Any]: + """ + Recommend priority markets for localization. + + Args: + current_market: Current/primary market + budget_level: 'low', 'medium', or 'high' + target_market_count: Number of markets to target + + Returns: + Prioritized market recommendations + """ + # Determine tier priorities based on budget + if budget_level == 'low': + priority_tiers = ['tier_1'] + max_markets = min(target_market_count, 3) + elif budget_level == 'medium': + priority_tiers = ['tier_1', 'tier_2'] + max_markets = min(target_market_count, 8) + else: # high budget + priority_tiers = ['tier_1', 'tier_2', 'tier_3'] + max_markets = target_market_count + + # Collect markets from priority tiers + recommended_markets = [] + for tier in priority_tiers: + for market in self.PRIORITY_MARKETS[tier]: + if market['language'] != current_market: + recommended_markets.append({ + **market, + 'tier': tier, + 'estimated_translation_cost': self._estimate_translation_cost( + market['language'] + ) + }) + + # Sort by revenue share and limit + recommended_markets.sort(key=lambda x: x['revenue_share'], reverse=True) + recommended_markets = recommended_markets[:max_markets] + + # Calculate potential ROI + total_potential_revenue_share = sum(m['revenue_share'] for m in recommended_markets) + + return { + 'recommended_markets': recommended_markets, + 'total_markets': len(recommended_markets), + 'estimated_total_revenue_lift': f"{total_potential_revenue_share*100:.1f}%", + 'estimated_cost': self._estimate_total_localization_cost(recommended_markets), + 'implementation_priority': self._prioritize_implementation(recommended_markets) + } + + def translate_metadata( + self, + source_metadata: Dict[str, str], + source_language: str, + target_language: str, + platform: str = 'apple' + ) -> Dict[str, Any]: + """ + Generate localized metadata with character limit considerations. + + Args: + source_metadata: Original metadata (title, description, etc.) + source_language: Source language code (e.g., 'en') + target_language: Target language code (e.g., 'es') + platform: 'apple' or 'google' + + Returns: + Localized metadata with character limit validation + """ + # Get character multiplier + target_lang_code = target_language.split('-')[0] + char_multiplier = self.CHAR_MULTIPLIERS.get(target_lang_code, 1.0) + + # Platform-specific limits + if platform == 'apple': + limits = {'title': 30, 'subtitle': 30, 'description': 4000, 'keywords': 100} + else: + limits = {'title': 50, 'short_description': 80, 'description': 4000} + + localized_metadata = {} + warnings = [] + + for field, text in source_metadata.items(): + if field not in limits: + continue + + # Estimate target length + estimated_length = int(len(text) * char_multiplier) + limit = limits[field] + + localized_metadata[field] = { + 'original_text': text, + 'original_length': len(text), + 'estimated_target_length': estimated_length, + 'character_limit': limit, + 'fits_within_limit': estimated_length <= limit, + 'translation_notes': self._get_translation_notes( + field, + target_language, + estimated_length, + limit + ) + } + + if estimated_length > limit: + warnings.append( + f"{field}: Estimated length ({estimated_length}) may exceed limit ({limit}) - " + f"condensing may be required" + ) + + return { + 'source_language': source_language, + 'target_language': target_language, + 'platform': platform, + 'localized_fields': localized_metadata, + 'character_multiplier': char_multiplier, + 'warnings': warnings, + 'recommendations': self._generate_translation_recommendations( + target_language, + warnings + ) + } + + def adapt_keywords( + self, + source_keywords: List[str], + source_language: str, + target_language: str, + target_market: str + ) -> Dict[str, Any]: + """ + Adapt keywords for target market (not just direct translation). + + Args: + source_keywords: Original keywords + source_language: Source language code + target_language: Target language code + target_market: Target market (e.g., 'France', 'Japan') + + Returns: + Adapted keyword recommendations + """ + # Cultural adaptation considerations + cultural_notes = self._get_cultural_keyword_considerations(target_market) + + # Search behavior differences + search_patterns = self._get_search_patterns(target_market) + + adapted_keywords = [] + for keyword in source_keywords: + adapted_keywords.append({ + 'source_keyword': keyword, + 'adaptation_strategy': self._determine_adaptation_strategy( + keyword, + target_market + ), + 'cultural_considerations': cultural_notes.get(keyword, []), + 'priority': 'high' if keyword in source_keywords[:3] else 'medium' + }) + + return { + 'source_language': source_language, + 'target_language': target_language, + 'target_market': target_market, + 'adapted_keywords': adapted_keywords, + 'search_behavior_notes': search_patterns, + 'recommendations': [ + 'Use native speakers for keyword research', + 'Test keywords with local users before finalizing', + 'Consider local competitors\' keyword strategies', + 'Monitor search trends in target market' + ] + } + + def validate_translations( + self, + translated_metadata: Dict[str, str], + target_language: str, + platform: str = 'apple' + ) -> Dict[str, Any]: + """ + Validate translated metadata for character limits and quality. + + Args: + translated_metadata: Translated text fields + target_language: Target language code + platform: 'apple' or 'google' + + Returns: + Validation report + """ + # Platform limits + if platform == 'apple': + limits = {'title': 30, 'subtitle': 30, 'description': 4000, 'keywords': 100} + else: + limits = {'title': 50, 'short_description': 80, 'description': 4000} + + validation_results = { + 'is_valid': True, + 'field_validations': {}, + 'errors': [], + 'warnings': [] + } + + for field, text in translated_metadata.items(): + if field not in limits: + continue + + actual_length = len(text) + limit = limits[field] + is_within_limit = actual_length <= limit + + validation_results['field_validations'][field] = { + 'text': text, + 'length': actual_length, + 'limit': limit, + 'is_valid': is_within_limit, + 'usage_percentage': round((actual_length / limit) * 100, 1) + } + + if not is_within_limit: + validation_results['is_valid'] = False + validation_results['errors'].append( + f"{field} exceeds limit: {actual_length}/{limit} characters" + ) + + # Quality checks + quality_issues = self._check_translation_quality( + translated_metadata, + target_language + ) + + validation_results['quality_checks'] = quality_issues + + if quality_issues: + validation_results['warnings'].extend( + [f"Quality issue: {issue}" for issue in quality_issues] + ) + + return validation_results + + def calculate_localization_roi( + self, + target_markets: List[str], + current_monthly_downloads: int, + localization_cost: float, + expected_lift_percentage: float = 0.15 + ) -> Dict[str, Any]: + """ + Estimate ROI of localization investment. + + Args: + target_markets: List of market codes + current_monthly_downloads: Current monthly downloads + localization_cost: Total cost to localize + expected_lift_percentage: Expected download increase (default 15%) + + Returns: + ROI analysis + """ + # Estimate market-specific lift + market_data = [] + total_expected_lift = 0 + + for market_code in target_markets: + # Find market in priority lists + market_info = None + for tier_name, markets in self.PRIORITY_MARKETS.items(): + for m in markets: + if m['language'] == market_code: + market_info = m + break + + if not market_info: + continue + + # Estimate downloads from this market + market_downloads = int(current_monthly_downloads * market_info['revenue_share']) + expected_increase = int(market_downloads * expected_lift_percentage) + total_expected_lift += expected_increase + + market_data.append({ + 'market': market_info['market'], + 'current_monthly_downloads': market_downloads, + 'expected_increase': expected_increase, + 'revenue_potential': market_info['revenue_share'] + }) + + # Calculate payback period (assuming $2 revenue per download) + revenue_per_download = 2.0 + monthly_additional_revenue = total_expected_lift * revenue_per_download + payback_months = (localization_cost / monthly_additional_revenue) if monthly_additional_revenue > 0 else float('inf') + + return { + 'markets_analyzed': len(market_data), + 'market_breakdown': market_data, + 'total_expected_monthly_lift': total_expected_lift, + 'expected_monthly_revenue_increase': f"${monthly_additional_revenue:,.2f}", + 'localization_cost': f"${localization_cost:,.2f}", + 'payback_period_months': round(payback_months, 1) if payback_months != float('inf') else 'N/A', + 'annual_roi': f"{((monthly_additional_revenue * 12 - localization_cost) / localization_cost * 100):.1f}%" if payback_months != float('inf') else 'Negative', + 'recommendation': self._generate_roi_recommendation(payback_months) + } + + def _estimate_translation_cost(self, language: str) -> Dict[str, float]: + """Estimate translation cost for a language.""" + # Base cost per word (professional translation) + base_cost_per_word = 0.12 + + # Language-specific multipliers + multipliers = { + 'zh-CN': 1.5, # Chinese requires specialist + 'ja-JP': 1.5, # Japanese requires specialist + 'ko-KR': 1.3, + 'ar-SA': 1.4, # Arabic (right-to-left) + 'default': 1.0 + } + + multiplier = multipliers.get(language, multipliers['default']) + + # Typical word counts for app store metadata + typical_word_counts = { + 'title': 5, + 'subtitle': 5, + 'description': 300, + 'keywords': 20, + 'screenshots': 50 # Caption text + } + + total_words = sum(typical_word_counts.values()) + estimated_cost = total_words * base_cost_per_word * multiplier + + return { + 'cost_per_word': base_cost_per_word * multiplier, + 'total_words': total_words, + 'estimated_cost': round(estimated_cost, 2) + } + + def _estimate_total_localization_cost(self, markets: List[Dict[str, Any]]) -> str: + """Estimate total cost for multiple markets.""" + total = sum(m['estimated_translation_cost']['estimated_cost'] for m in markets) + return f"${total:,.2f}" + + def _prioritize_implementation(self, markets: List[Dict[str, Any]]) -> List[Dict[str, str]]: + """Create phased implementation plan.""" + phases = [] + + # Phase 1: Top revenue markets + phase_1 = [m for m in markets[:3]] + if phase_1: + phases.append({ + 'phase': 'Phase 1 (First 30 days)', + 'markets': ', '.join([m['market'] for m in phase_1]), + 'rationale': 'Highest revenue potential markets' + }) + + # Phase 2: Remaining tier 1 and top tier 2 + phase_2 = [m for m in markets[3:6]] + if phase_2: + phases.append({ + 'phase': 'Phase 2 (Days 31-60)', + 'markets': ', '.join([m['market'] for m in phase_2]), + 'rationale': 'Strong revenue markets with good ROI' + }) + + # Phase 3: Remaining markets + phase_3 = [m for m in markets[6:]] + if phase_3: + phases.append({ + 'phase': 'Phase 3 (Days 61-90)', + 'markets': ', '.join([m['market'] for m in phase_3]), + 'rationale': 'Complete global coverage' + }) + + return phases + + def _get_translation_notes( + self, + field: str, + target_language: str, + estimated_length: int, + limit: int + ) -> List[str]: + """Get translation-specific notes for field.""" + notes = [] + + if estimated_length > limit: + notes.append(f"Condensing required - aim for {limit - 10} characters to allow buffer") + + if field == 'title' and target_language.startswith('zh'): + notes.append("Chinese characters convey more meaning - may need fewer characters") + + if field == 'keywords' and target_language.startswith('de'): + notes.append("German compound words may be longer - prioritize shorter keywords") + + return notes + + def _generate_translation_recommendations( + self, + target_language: str, + warnings: List[str] + ) -> List[str]: + """Generate translation recommendations.""" + recommendations = [ + "Use professional native speakers for translation", + "Test translations with local users before finalizing" + ] + + if warnings: + recommendations.append("Work with translator to condense text while preserving meaning") + + if target_language.startswith('zh') or target_language.startswith('ja'): + recommendations.append("Consider cultural context and local idioms") + + return recommendations + + def _get_cultural_keyword_considerations(self, target_market: str) -> Dict[str, List[str]]: + """Get cultural considerations for keywords by market.""" + # Simplified example - real implementation would be more comprehensive + considerations = { + 'China': ['Avoid politically sensitive terms', 'Consider local alternatives to blocked services'], + 'Japan': ['Honorific language important', 'Technical terms often use katakana'], + 'Germany': ['Privacy and security terms resonate', 'Efficiency and quality valued'], + 'France': ['French language protection laws', 'Prefer French terms over English'], + 'default': ['Research local search behavior', 'Test with native speakers'] + } + + return considerations.get(target_market, considerations['default']) + + def _get_search_patterns(self, target_market: str) -> List[str]: + """Get search pattern notes for market.""" + patterns = { + 'China': ['Use both simplified characters and romanization', 'Brand names often romanized'], + 'Japan': ['Mix of kanji, hiragana, and katakana', 'English words common in tech'], + 'Germany': ['Compound words common', 'Specific technical terminology'], + 'default': ['Research local search trends', 'Monitor competitor keywords'] + } + + return patterns.get(target_market, patterns['default']) + + def _determine_adaptation_strategy(self, keyword: str, target_market: str) -> str: + """Determine how to adapt keyword for market.""" + # Simplified logic + if target_market in ['China', 'Japan', 'Korea']: + return 'full_localization' # Complete translation needed + elif target_market in ['Germany', 'France', 'Spain']: + return 'adapt_and_translate' # Some adaptation needed + else: + return 'direct_translation' # Direct translation usually sufficient + + def _check_translation_quality( + self, + translated_metadata: Dict[str, str], + target_language: str + ) -> List[str]: + """Basic quality checks for translations.""" + issues = [] + + # Check for untranslated placeholders + for field, text in translated_metadata.items(): + if '[' in text or '{' in text or 'TODO' in text.upper(): + issues.append(f"{field} contains placeholder text") + + # Check for excessive punctuation + for field, text in translated_metadata.items(): + if text.count('!') > 3: + issues.append(f"{field} has excessive exclamation marks") + + return issues + + def _generate_roi_recommendation(self, payback_months: float) -> str: + """Generate ROI recommendation.""" + if payback_months <= 3: + return "Excellent ROI - proceed immediately" + elif payback_months <= 6: + return "Good ROI - recommended investment" + elif payback_months <= 12: + return "Moderate ROI - consider if strategic market" + else: + return "Low ROI - reconsider or focus on higher-priority markets first" + + +def plan_localization_strategy( + current_market: str, + budget_level: str, + monthly_downloads: int +) -> Dict[str, Any]: + """ + Convenience function to plan localization strategy. + + Args: + current_market: Current market code + budget_level: Budget level + monthly_downloads: Current monthly downloads + + Returns: + Complete localization plan + """ + helper = LocalizationHelper() + + target_markets = helper.identify_target_markets( + current_market=current_market, + budget_level=budget_level + ) + + # Extract market codes + market_codes = [m['language'] for m in target_markets['recommended_markets']] + + # Calculate ROI + estimated_cost = float(target_markets['estimated_cost'].replace('$', '').replace(',', '')) + + roi_analysis = helper.calculate_localization_roi( + market_codes, + monthly_downloads, + estimated_cost + ) + + return { + 'target_markets': target_markets, + 'roi_analysis': roi_analysis + } diff --git a/skills/app-store-optimization/scripts/metadata_optimizer.py b/skills/app-store-optimization/scripts/metadata_optimizer.py new file mode 100644 index 00000000..7b506140 --- /dev/null +++ b/skills/app-store-optimization/scripts/metadata_optimizer.py @@ -0,0 +1,581 @@ +""" +Metadata optimization module for App Store Optimization. +Optimizes titles, descriptions, and keyword fields with platform-specific character limit validation. +""" + +from typing import Dict, List, Any, Optional, Tuple +import re + + +class MetadataOptimizer: + """Optimizes app store metadata for maximum discoverability and conversion.""" + + # Platform-specific character limits + CHAR_LIMITS = { + 'apple': { + 'title': 30, + 'subtitle': 30, + 'promotional_text': 170, + 'description': 4000, + 'keywords': 100, + 'whats_new': 4000 + }, + 'google': { + 'title': 50, + 'short_description': 80, + 'full_description': 4000 + } + } + + def __init__(self, platform: str = 'apple'): + """ + Initialize metadata optimizer. + + Args: + platform: 'apple' or 'google' + """ + if platform not in ['apple', 'google']: + raise ValueError("Platform must be 'apple' or 'google'") + + self.platform = platform + self.limits = self.CHAR_LIMITS[platform] + + def optimize_title( + self, + app_name: str, + target_keywords: List[str], + include_brand: bool = True + ) -> Dict[str, Any]: + """ + Optimize app title with keyword integration. + + Args: + app_name: Your app's brand name + target_keywords: List of keywords to potentially include + include_brand: Whether to include brand name + + Returns: + Optimized title options with analysis + """ + max_length = self.limits['title'] + + title_options = [] + + # Option 1: Brand name only + if include_brand: + option1 = app_name[:max_length] + title_options.append({ + 'title': option1, + 'length': len(option1), + 'remaining_chars': max_length - len(option1), + 'keywords_included': [], + 'strategy': 'brand_only', + 'pros': ['Maximum brand recognition', 'Clean and simple'], + 'cons': ['No keyword targeting', 'Lower discoverability'] + }) + + # Option 2: Brand + Primary Keyword + if target_keywords: + primary_keyword = target_keywords[0] + option2 = self._build_title_with_keywords( + app_name, + [primary_keyword], + max_length + ) + if option2: + title_options.append({ + 'title': option2, + 'length': len(option2), + 'remaining_chars': max_length - len(option2), + 'keywords_included': [primary_keyword], + 'strategy': 'brand_plus_primary', + 'pros': ['Targets main keyword', 'Maintains brand identity'], + 'cons': ['Limited keyword coverage'] + }) + + # Option 3: Brand + Multiple Keywords (if space allows) + if len(target_keywords) > 1: + option3 = self._build_title_with_keywords( + app_name, + target_keywords[:2], + max_length + ) + if option3: + title_options.append({ + 'title': option3, + 'length': len(option3), + 'remaining_chars': max_length - len(option3), + 'keywords_included': target_keywords[:2], + 'strategy': 'brand_plus_multiple', + 'pros': ['Multiple keyword targets', 'Better discoverability'], + 'cons': ['May feel cluttered', 'Less brand focus'] + }) + + # Option 4: Keyword-first approach (for new apps) + if target_keywords and not include_brand: + option4 = " ".join(target_keywords[:2])[:max_length] + title_options.append({ + 'title': option4, + 'length': len(option4), + 'remaining_chars': max_length - len(option4), + 'keywords_included': target_keywords[:2], + 'strategy': 'keyword_first', + 'pros': ['Maximum SEO benefit', 'Clear functionality'], + 'cons': ['No brand recognition', 'Generic appearance'] + }) + + return { + 'platform': self.platform, + 'max_length': max_length, + 'options': title_options, + 'recommendation': self._recommend_title_option(title_options) + } + + def optimize_description( + self, + app_info: Dict[str, Any], + target_keywords: List[str], + description_type: str = 'full' + ) -> Dict[str, Any]: + """ + Optimize app description with keyword integration and conversion focus. + + Args: + app_info: Dict with 'name', 'key_features', 'unique_value', 'target_audience' + target_keywords: List of keywords to integrate naturally + description_type: 'full', 'short' (Google), 'subtitle' (Apple) + + Returns: + Optimized description with analysis + """ + if description_type == 'short' and self.platform == 'google': + return self._optimize_short_description(app_info, target_keywords) + elif description_type == 'subtitle' and self.platform == 'apple': + return self._optimize_subtitle(app_info, target_keywords) + else: + return self._optimize_full_description(app_info, target_keywords) + + def optimize_keyword_field( + self, + target_keywords: List[str], + app_title: str = "", + app_description: str = "" + ) -> Dict[str, Any]: + """ + Optimize Apple's 100-character keyword field. + + Rules: + - No spaces between commas + - No plural forms if singular exists + - No duplicates + - Keywords in title/subtitle are already indexed + + Args: + target_keywords: List of target keywords + app_title: Current app title (to avoid duplication) + app_description: Current description (to check coverage) + + Returns: + Optimized keyword field (comma-separated, no spaces) + """ + if self.platform != 'apple': + return {'error': 'Keyword field optimization only applies to Apple App Store'} + + max_length = self.limits['keywords'] + + # Extract words already in title (these don't need to be in keyword field) + title_words = set(app_title.lower().split()) if app_title else set() + + # Process keywords + processed_keywords = [] + for keyword in target_keywords: + keyword_lower = keyword.lower().strip() + + # Skip if already in title + if keyword_lower in title_words: + continue + + # Remove duplicates and process + words = keyword_lower.split() + for word in words: + if word not in processed_keywords and word not in title_words: + processed_keywords.append(word) + + # Remove plurals if singular exists + deduplicated = self._remove_plural_duplicates(processed_keywords) + + # Build keyword field within 100 character limit + keyword_field = self._build_keyword_field(deduplicated, max_length) + + # Calculate keyword density in description + density = self._calculate_coverage(target_keywords, app_description) + + return { + 'keyword_field': keyword_field, + 'length': len(keyword_field), + 'remaining_chars': max_length - len(keyword_field), + 'keywords_included': keyword_field.split(','), + 'keywords_count': len(keyword_field.split(',')), + 'keywords_excluded': [kw for kw in target_keywords if kw.lower() not in keyword_field], + 'description_coverage': density, + 'optimization_tips': [ + 'Keywords in title are auto-indexed - no need to repeat', + 'Use singular forms only (Apple indexes plurals automatically)', + 'No spaces between commas to maximize character usage', + 'Update keyword field with each app update to test variations' + ] + } + + def validate_character_limits( + self, + metadata: Dict[str, str] + ) -> Dict[str, Any]: + """ + Validate all metadata fields against platform character limits. + + Args: + metadata: Dictionary of field_name: value + + Returns: + Validation report with errors and warnings + """ + validation_results = { + 'is_valid': True, + 'errors': [], + 'warnings': [], + 'field_status': {} + } + + for field_name, value in metadata.items(): + if field_name not in self.limits: + validation_results['warnings'].append( + f"Unknown field '{field_name}' for {self.platform} platform" + ) + continue + + max_length = self.limits[field_name] + actual_length = len(value) + remaining = max_length - actual_length + + field_status = { + 'value': value, + 'length': actual_length, + 'limit': max_length, + 'remaining': remaining, + 'is_valid': actual_length <= max_length, + 'usage_percentage': round((actual_length / max_length) * 100, 1) + } + + validation_results['field_status'][field_name] = field_status + + if actual_length > max_length: + validation_results['is_valid'] = False + validation_results['errors'].append( + f"'{field_name}' exceeds limit: {actual_length}/{max_length} chars" + ) + elif remaining > max_length * 0.2: # More than 20% unused + validation_results['warnings'].append( + f"'{field_name}' under-utilizes space: {remaining} chars remaining" + ) + + return validation_results + + def calculate_keyword_density( + self, + text: str, + target_keywords: List[str] + ) -> Dict[str, Any]: + """ + Calculate keyword density in text. + + Args: + text: Text to analyze + target_keywords: Keywords to check + + Returns: + Density analysis + """ + text_lower = text.lower() + total_words = len(text_lower.split()) + + keyword_densities = {} + for keyword in target_keywords: + keyword_lower = keyword.lower() + count = text_lower.count(keyword_lower) + density = (count / total_words * 100) if total_words > 0 else 0 + + keyword_densities[keyword] = { + 'occurrences': count, + 'density_percentage': round(density, 2), + 'status': self._assess_density(density) + } + + # Overall assessment + total_keyword_occurrences = sum(kw['occurrences'] for kw in keyword_densities.values()) + overall_density = (total_keyword_occurrences / total_words * 100) if total_words > 0 else 0 + + return { + 'total_words': total_words, + 'keyword_densities': keyword_densities, + 'overall_keyword_density': round(overall_density, 2), + 'assessment': self._assess_overall_density(overall_density), + 'recommendations': self._generate_density_recommendations(keyword_densities) + } + + def _build_title_with_keywords( + self, + app_name: str, + keywords: List[str], + max_length: int + ) -> Optional[str]: + """Build title combining app name and keywords within limit.""" + separators = [' - ', ': ', ' | '] + + for sep in separators: + for kw in keywords: + title = f"{app_name}{sep}{kw}" + if len(title) <= max_length: + return title + + return None + + def _optimize_short_description( + self, + app_info: Dict[str, Any], + target_keywords: List[str] + ) -> Dict[str, Any]: + """Optimize Google Play short description (80 chars).""" + max_length = self.limits['short_description'] + + # Focus on unique value proposition with primary keyword + unique_value = app_info.get('unique_value', '') + primary_keyword = target_keywords[0] if target_keywords else '' + + # Template: [Primary Keyword] - [Unique Value] + short_desc = f"{primary_keyword.title()} - {unique_value}"[:max_length] + + return { + 'short_description': short_desc, + 'length': len(short_desc), + 'remaining_chars': max_length - len(short_desc), + 'keywords_included': [primary_keyword] if primary_keyword in short_desc.lower() else [], + 'strategy': 'keyword_value_proposition' + } + + def _optimize_subtitle( + self, + app_info: Dict[str, Any], + target_keywords: List[str] + ) -> Dict[str, Any]: + """Optimize Apple App Store subtitle (30 chars).""" + max_length = self.limits['subtitle'] + + # Very concise - primary keyword or key feature + primary_keyword = target_keywords[0] if target_keywords else '' + key_feature = app_info.get('key_features', [''])[0] if app_info.get('key_features') else '' + + options = [ + primary_keyword[:max_length], + key_feature[:max_length], + f"{primary_keyword} App"[:max_length] + ] + + return { + 'subtitle_options': [opt for opt in options if opt], + 'max_length': max_length, + 'recommendation': options[0] if options else '' + } + + def _optimize_full_description( + self, + app_info: Dict[str, Any], + target_keywords: List[str] + ) -> Dict[str, Any]: + """Optimize full app description (4000 chars for both platforms).""" + max_length = self.limits.get('description', self.limits.get('full_description', 4000)) + + # Structure: Hook → Features → Benefits → Social Proof → CTA + sections = [] + + # Hook (with primary keyword) + primary_keyword = target_keywords[0] if target_keywords else '' + unique_value = app_info.get('unique_value', '') + hook = f"{unique_value} {primary_keyword.title()} that helps you achieve more.\n\n" + sections.append(hook) + + # Features (with keywords naturally integrated) + features = app_info.get('key_features', []) + if features: + sections.append("KEY FEATURES:\n") + for i, feature in enumerate(features[:5], 1): + # Integrate keywords naturally + feature_text = f"• {feature}" + if i <= len(target_keywords): + keyword = target_keywords[i-1] + if keyword.lower() not in feature.lower(): + feature_text = f"• {feature} with {keyword}" + sections.append(f"{feature_text}\n") + sections.append("\n") + + # Benefits + target_audience = app_info.get('target_audience', 'users') + sections.append(f"PERFECT FOR:\n{target_audience}\n\n") + + # Social proof placeholder + sections.append("WHY USERS LOVE US:\n") + sections.append("Join thousands of satisfied users who have transformed their workflow.\n\n") + + # CTA + sections.append("Download now and start experiencing the difference!") + + # Combine and validate length + full_description = "".join(sections) + if len(full_description) > max_length: + full_description = full_description[:max_length-3] + "..." + + # Calculate keyword density + density = self.calculate_keyword_density(full_description, target_keywords) + + return { + 'full_description': full_description, + 'length': len(full_description), + 'remaining_chars': max_length - len(full_description), + 'keyword_analysis': density, + 'structure': { + 'has_hook': True, + 'has_features': len(features) > 0, + 'has_benefits': True, + 'has_cta': True + } + } + + def _remove_plural_duplicates(self, keywords: List[str]) -> List[str]: + """Remove plural forms if singular exists.""" + deduplicated = [] + singular_set = set() + + for keyword in keywords: + if keyword.endswith('s') and len(keyword) > 1: + singular = keyword[:-1] + if singular not in singular_set: + deduplicated.append(singular) + singular_set.add(singular) + else: + if keyword not in singular_set: + deduplicated.append(keyword) + singular_set.add(keyword) + + return deduplicated + + def _build_keyword_field(self, keywords: List[str], max_length: int) -> str: + """Build comma-separated keyword field within character limit.""" + keyword_field = "" + + for keyword in keywords: + test_field = f"{keyword_field},{keyword}" if keyword_field else keyword + if len(test_field) <= max_length: + keyword_field = test_field + else: + break + + return keyword_field + + def _calculate_coverage(self, keywords: List[str], text: str) -> Dict[str, int]: + """Calculate how many keywords are covered in text.""" + text_lower = text.lower() + coverage = {} + + for keyword in keywords: + coverage[keyword] = text_lower.count(keyword.lower()) + + return coverage + + def _assess_density(self, density: float) -> str: + """Assess individual keyword density.""" + if density < 0.5: + return "too_low" + elif density <= 2.5: + return "optimal" + else: + return "too_high" + + def _assess_overall_density(self, density: float) -> str: + """Assess overall keyword density.""" + if density < 2: + return "Under-optimized: Consider adding more keyword variations" + elif density <= 5: + return "Optimal: Good keyword integration without stuffing" + elif density <= 8: + return "High: Approaching keyword stuffing - reduce keyword usage" + else: + return "Too High: Keyword stuffing detected - rewrite for natural flow" + + def _generate_density_recommendations( + self, + keyword_densities: Dict[str, Dict[str, Any]] + ) -> List[str]: + """Generate recommendations based on keyword density analysis.""" + recommendations = [] + + for keyword, data in keyword_densities.items(): + if data['status'] == 'too_low': + recommendations.append( + f"Increase usage of '{keyword}' - currently only {data['occurrences']} times" + ) + elif data['status'] == 'too_high': + recommendations.append( + f"Reduce usage of '{keyword}' - appears {data['occurrences']} times (keyword stuffing risk)" + ) + + if not recommendations: + recommendations.append("Keyword density is well-balanced") + + return recommendations + + def _recommend_title_option(self, options: List[Dict[str, Any]]) -> str: + """Recommend best title option based on strategy.""" + if not options: + return "No valid options available" + + # Prefer brand_plus_primary for established apps + for option in options: + if option['strategy'] == 'brand_plus_primary': + return f"Recommended: '{option['title']}' (Balance of brand and SEO)" + + # Fallback to first option + return f"Recommended: '{options[0]['title']}' ({options[0]['strategy']})" + + +def optimize_app_metadata( + platform: str, + app_info: Dict[str, Any], + target_keywords: List[str] +) -> Dict[str, Any]: + """ + Convenience function to optimize all metadata fields. + + Args: + platform: 'apple' or 'google' + app_info: App information dictionary + target_keywords: Target keywords list + + Returns: + Complete metadata optimization package + """ + optimizer = MetadataOptimizer(platform) + + return { + 'platform': platform, + 'title': optimizer.optimize_title( + app_info['name'], + target_keywords + ), + 'description': optimizer.optimize_description( + app_info, + target_keywords, + 'full' + ), + 'keyword_field': optimizer.optimize_keyword_field( + target_keywords + ) if platform == 'apple' else None + } diff --git a/skills/app-store-optimization/scripts/review_analyzer.py b/skills/app-store-optimization/scripts/review_analyzer.py new file mode 100644 index 00000000..4ce124d5 --- /dev/null +++ b/skills/app-store-optimization/scripts/review_analyzer.py @@ -0,0 +1,714 @@ +""" +Review analysis module for App Store Optimization. +Analyzes user reviews for sentiment, issues, and feature requests. +""" + +from typing import Dict, List, Any, Optional, Tuple +from collections import Counter +import re + + +class ReviewAnalyzer: + """Analyzes user reviews for actionable insights.""" + + # Sentiment keywords + POSITIVE_KEYWORDS = [ + 'great', 'awesome', 'excellent', 'amazing', 'love', 'best', 'perfect', + 'fantastic', 'wonderful', 'brilliant', 'outstanding', 'superb' + ] + + NEGATIVE_KEYWORDS = [ + 'bad', 'terrible', 'awful', 'horrible', 'hate', 'worst', 'useless', + 'broken', 'crash', 'bug', 'slow', 'disappointing', 'frustrating' + ] + + # Issue indicators + ISSUE_KEYWORDS = [ + 'crash', 'bug', 'error', 'broken', 'not working', 'doesnt work', + 'freezes', 'slow', 'laggy', 'glitch', 'problem', 'issue', 'fail' + ] + + # Feature request indicators + FEATURE_REQUEST_KEYWORDS = [ + 'wish', 'would be nice', 'should add', 'need', 'want', 'hope', + 'please add', 'missing', 'lacks', 'feature request' + ] + + def __init__(self, app_name: str): + """ + Initialize review analyzer. + + Args: + app_name: Name of the app + """ + self.app_name = app_name + self.reviews = [] + self.analysis_cache = {} + + def analyze_sentiment( + self, + reviews: List[Dict[str, Any]] + ) -> Dict[str, Any]: + """ + Analyze sentiment across reviews. + + Args: + reviews: List of review dicts with 'text', 'rating', 'date' + + Returns: + Sentiment analysis summary + """ + self.reviews = reviews + + sentiment_counts = { + 'positive': 0, + 'neutral': 0, + 'negative': 0 + } + + detailed_sentiments = [] + + for review in reviews: + text = review.get('text', '').lower() + rating = review.get('rating', 3) + + # Calculate sentiment score + sentiment_score = self._calculate_sentiment_score(text, rating) + sentiment_category = self._categorize_sentiment(sentiment_score) + + sentiment_counts[sentiment_category] += 1 + + detailed_sentiments.append({ + 'review_id': review.get('id', ''), + 'rating': rating, + 'sentiment_score': sentiment_score, + 'sentiment': sentiment_category, + 'text_preview': text[:100] + '...' if len(text) > 100 else text + }) + + # Calculate percentages + total = len(reviews) + sentiment_distribution = { + 'positive': round((sentiment_counts['positive'] / total) * 100, 1) if total > 0 else 0, + 'neutral': round((sentiment_counts['neutral'] / total) * 100, 1) if total > 0 else 0, + 'negative': round((sentiment_counts['negative'] / total) * 100, 1) if total > 0 else 0 + } + + # Calculate average rating + avg_rating = sum(r.get('rating', 0) for r in reviews) / total if total > 0 else 0 + + return { + 'total_reviews_analyzed': total, + 'average_rating': round(avg_rating, 2), + 'sentiment_distribution': sentiment_distribution, + 'sentiment_counts': sentiment_counts, + 'sentiment_trend': self._assess_sentiment_trend(sentiment_distribution), + 'detailed_sentiments': detailed_sentiments[:50] # Limit output + } + + def extract_common_themes( + self, + reviews: List[Dict[str, Any]], + min_mentions: int = 3 + ) -> Dict[str, Any]: + """ + Extract frequently mentioned themes and topics. + + Args: + reviews: List of review dicts + min_mentions: Minimum mentions to be considered common + + Returns: + Common themes analysis + """ + # Extract all words from reviews + all_words = [] + all_phrases = [] + + for review in reviews: + text = review.get('text', '').lower() + # Clean text + text = re.sub(r'[^\w\s]', ' ', text) + words = text.split() + + # Filter out common words + stop_words = { + 'the', 'and', 'for', 'with', 'this', 'that', 'from', 'have', + 'app', 'apps', 'very', 'really', 'just', 'but', 'not', 'you' + } + words = [w for w in words if w not in stop_words and len(w) > 3] + + all_words.extend(words) + + # Extract 2-3 word phrases + for i in range(len(words) - 1): + phrase = f"{words[i]} {words[i+1]}" + all_phrases.append(phrase) + + # Count frequency + word_freq = Counter(all_words) + phrase_freq = Counter(all_phrases) + + # Filter by min_mentions + common_words = [ + {'word': word, 'mentions': count} + for word, count in word_freq.most_common(30) + if count >= min_mentions + ] + + common_phrases = [ + {'phrase': phrase, 'mentions': count} + for phrase, count in phrase_freq.most_common(20) + if count >= min_mentions + ] + + # Categorize themes + themes = self._categorize_themes(common_words, common_phrases) + + return { + 'common_words': common_words, + 'common_phrases': common_phrases, + 'identified_themes': themes, + 'insights': self._generate_theme_insights(themes) + } + + def identify_issues( + self, + reviews: List[Dict[str, Any]], + rating_threshold: int = 3 + ) -> Dict[str, Any]: + """ + Identify bugs, crashes, and other issues from reviews. + + Args: + reviews: List of review dicts + rating_threshold: Only analyze reviews at or below this rating + + Returns: + Issue identification report + """ + issues = [] + + for review in reviews: + rating = review.get('rating', 5) + if rating > rating_threshold: + continue + + text = review.get('text', '').lower() + + # Check for issue keywords + mentioned_issues = [] + for keyword in self.ISSUE_KEYWORDS: + if keyword in text: + mentioned_issues.append(keyword) + + if mentioned_issues: + issues.append({ + 'review_id': review.get('id', ''), + 'rating': rating, + 'date': review.get('date', ''), + 'issue_keywords': mentioned_issues, + 'text': text[:200] + '...' if len(text) > 200 else text + }) + + # Group by issue type + issue_frequency = Counter() + for issue in issues: + for keyword in issue['issue_keywords']: + issue_frequency[keyword] += 1 + + # Categorize issues + categorized_issues = self._categorize_issues(issues) + + # Calculate issue severity + severity_scores = self._calculate_issue_severity( + categorized_issues, + len(reviews) + ) + + return { + 'total_issues_found': len(issues), + 'issue_frequency': dict(issue_frequency.most_common(15)), + 'categorized_issues': categorized_issues, + 'severity_scores': severity_scores, + 'top_issues': self._rank_issues_by_severity(severity_scores), + 'recommendations': self._generate_issue_recommendations( + categorized_issues, + severity_scores + ) + } + + def find_feature_requests( + self, + reviews: List[Dict[str, Any]] + ) -> Dict[str, Any]: + """ + Extract feature requests and desired improvements. + + Args: + reviews: List of review dicts + + Returns: + Feature request analysis + """ + feature_requests = [] + + for review in reviews: + text = review.get('text', '').lower() + rating = review.get('rating', 3) + + # Check for feature request indicators + is_feature_request = any( + keyword in text + for keyword in self.FEATURE_REQUEST_KEYWORDS + ) + + if is_feature_request: + # Extract the specific request + request_text = self._extract_feature_request_text(text) + + feature_requests.append({ + 'review_id': review.get('id', ''), + 'rating': rating, + 'date': review.get('date', ''), + 'request_text': request_text, + 'full_review': text[:200] + '...' if len(text) > 200 else text + }) + + # Cluster similar requests + clustered_requests = self._cluster_feature_requests(feature_requests) + + # Prioritize based on frequency and rating context + prioritized_requests = self._prioritize_feature_requests(clustered_requests) + + return { + 'total_feature_requests': len(feature_requests), + 'clustered_requests': clustered_requests, + 'prioritized_requests': prioritized_requests, + 'implementation_recommendations': self._generate_feature_recommendations( + prioritized_requests + ) + } + + def track_sentiment_trends( + self, + reviews_by_period: Dict[str, List[Dict[str, Any]]] + ) -> Dict[str, Any]: + """ + Track sentiment changes over time. + + Args: + reviews_by_period: Dict of period_name: reviews + + Returns: + Trend analysis + """ + trends = [] + + for period, reviews in reviews_by_period.items(): + sentiment = self.analyze_sentiment(reviews) + + trends.append({ + 'period': period, + 'total_reviews': len(reviews), + 'average_rating': sentiment['average_rating'], + 'positive_percentage': sentiment['sentiment_distribution']['positive'], + 'negative_percentage': sentiment['sentiment_distribution']['negative'] + }) + + # Calculate trend direction + if len(trends) >= 2: + first_period = trends[0] + last_period = trends[-1] + + rating_change = last_period['average_rating'] - first_period['average_rating'] + sentiment_change = last_period['positive_percentage'] - first_period['positive_percentage'] + + trend_direction = self._determine_trend_direction( + rating_change, + sentiment_change + ) + else: + trend_direction = 'insufficient_data' + + return { + 'periods_analyzed': len(trends), + 'trend_data': trends, + 'trend_direction': trend_direction, + 'insights': self._generate_trend_insights(trends, trend_direction) + } + + def generate_response_templates( + self, + issue_category: str + ) -> List[Dict[str, str]]: + """ + Generate response templates for common review scenarios. + + Args: + issue_category: Category of issue ('crash', 'feature_request', 'positive', etc.) + + Returns: + Response templates + """ + templates = { + 'crash': [ + { + 'scenario': 'App crash reported', + 'template': "Thank you for bringing this to our attention. We're sorry you experienced a crash. " + "Our team is investigating this issue. Could you please share more details about when " + "this occurred (device model, iOS/Android version) by contacting support@[company].com? " + "We're committed to fixing this quickly." + }, + { + 'scenario': 'Crash already fixed', + 'template': "Thank you for your feedback. We've identified and fixed this crash issue in version [X.X]. " + "Please update to the latest version. If the problem persists, please reach out to " + "support@[company].com and we'll help you directly." + } + ], + 'bug': [ + { + 'scenario': 'Bug reported', + 'template': "Thanks for reporting this bug. We take these issues seriously. Our team is looking into it " + "and we'll have a fix in an upcoming update. We appreciate your patience and will notify you " + "when it's resolved." + } + ], + 'feature_request': [ + { + 'scenario': 'Feature request received', + 'template': "Thank you for this suggestion! We're always looking to improve [app_name]. We've added your " + "request to our roadmap and will consider it for a future update. Follow us @[social] for " + "updates on new features." + }, + { + 'scenario': 'Feature already planned', + 'template': "Great news! This feature is already on our roadmap and we're working on it. Stay tuned for " + "updates in the coming months. Thanks for your feedback!" + } + ], + 'positive': [ + { + 'scenario': 'Positive review', + 'template': "Thank you so much for your kind words! We're thrilled that you're enjoying [app_name]. " + "Reviews like yours motivate our team to keep improving. If you ever have suggestions, " + "we'd love to hear them!" + } + ], + 'negative_general': [ + { + 'scenario': 'General complaint', + 'template': "We're sorry to hear you're not satisfied with your experience. We'd like to make this right. " + "Please contact us at support@[company].com so we can understand the issue better and help " + "you directly. Thank you for giving us a chance to improve." + } + ] + } + + return templates.get(issue_category, templates['negative_general']) + + def _calculate_sentiment_score(self, text: str, rating: int) -> float: + """Calculate sentiment score (-1 to 1).""" + # Start with rating-based score + rating_score = (rating - 3) / 2 # Convert 1-5 to -1 to 1 + + # Adjust based on text sentiment + positive_count = sum(1 for keyword in self.POSITIVE_KEYWORDS if keyword in text) + negative_count = sum(1 for keyword in self.NEGATIVE_KEYWORDS if keyword in text) + + text_score = (positive_count - negative_count) / 10 # Normalize + + # Weighted average (60% rating, 40% text) + final_score = (rating_score * 0.6) + (text_score * 0.4) + + return max(min(final_score, 1.0), -1.0) + + def _categorize_sentiment(self, score: float) -> str: + """Categorize sentiment score.""" + if score > 0.3: + return 'positive' + elif score < -0.3: + return 'negative' + else: + return 'neutral' + + def _assess_sentiment_trend(self, distribution: Dict[str, float]) -> str: + """Assess overall sentiment trend.""" + positive = distribution['positive'] + negative = distribution['negative'] + + if positive > 70: + return 'very_positive' + elif positive > 50: + return 'positive' + elif negative > 30: + return 'concerning' + elif negative > 50: + return 'critical' + else: + return 'mixed' + + def _categorize_themes( + self, + common_words: List[Dict[str, Any]], + common_phrases: List[Dict[str, Any]] + ) -> Dict[str, List[str]]: + """Categorize themes from words and phrases.""" + themes = { + 'features': [], + 'performance': [], + 'usability': [], + 'support': [], + 'pricing': [] + } + + # Keywords for each category + feature_keywords = {'feature', 'functionality', 'option', 'tool'} + performance_keywords = {'fast', 'slow', 'crash', 'lag', 'speed', 'performance'} + usability_keywords = {'easy', 'difficult', 'intuitive', 'confusing', 'interface', 'design'} + support_keywords = {'support', 'help', 'customer', 'service', 'response'} + pricing_keywords = {'price', 'cost', 'expensive', 'cheap', 'subscription', 'free'} + + for word_data in common_words: + word = word_data['word'] + if any(kw in word for kw in feature_keywords): + themes['features'].append(word) + elif any(kw in word for kw in performance_keywords): + themes['performance'].append(word) + elif any(kw in word for kw in usability_keywords): + themes['usability'].append(word) + elif any(kw in word for kw in support_keywords): + themes['support'].append(word) + elif any(kw in word for kw in pricing_keywords): + themes['pricing'].append(word) + + return {k: v for k, v in themes.items() if v} # Remove empty categories + + def _generate_theme_insights(self, themes: Dict[str, List[str]]) -> List[str]: + """Generate insights from themes.""" + insights = [] + + for category, keywords in themes.items(): + if keywords: + insights.append( + f"{category.title()}: Users frequently mention {', '.join(keywords[:3])}" + ) + + return insights[:5] + + def _categorize_issues(self, issues: List[Dict[str, Any]]) -> Dict[str, List[Dict[str, Any]]]: + """Categorize issues by type.""" + categories = { + 'crashes': [], + 'bugs': [], + 'performance': [], + 'compatibility': [] + } + + for issue in issues: + keywords = issue['issue_keywords'] + + if 'crash' in keywords or 'freezes' in keywords: + categories['crashes'].append(issue) + elif 'bug' in keywords or 'error' in keywords or 'broken' in keywords: + categories['bugs'].append(issue) + elif 'slow' in keywords or 'laggy' in keywords: + categories['performance'].append(issue) + else: + categories['compatibility'].append(issue) + + return {k: v for k, v in categories.items() if v} + + def _calculate_issue_severity( + self, + categorized_issues: Dict[str, List[Dict[str, Any]]], + total_reviews: int + ) -> Dict[str, Dict[str, Any]]: + """Calculate severity scores for each issue category.""" + severity_scores = {} + + for category, issues in categorized_issues.items(): + count = len(issues) + percentage = (count / total_reviews) * 100 if total_reviews > 0 else 0 + + # Calculate average rating of affected reviews + avg_rating = sum(i['rating'] for i in issues) / count if count > 0 else 0 + + # Severity score (0-100) + severity = min((percentage * 10) + ((5 - avg_rating) * 10), 100) + + severity_scores[category] = { + 'count': count, + 'percentage': round(percentage, 2), + 'average_rating': round(avg_rating, 2), + 'severity_score': round(severity, 1), + 'priority': 'critical' if severity > 70 else ('high' if severity > 40 else 'medium') + } + + return severity_scores + + def _rank_issues_by_severity( + self, + severity_scores: Dict[str, Dict[str, Any]] + ) -> List[Dict[str, Any]]: + """Rank issues by severity score.""" + ranked = sorted( + [{'category': cat, **data} for cat, data in severity_scores.items()], + key=lambda x: x['severity_score'], + reverse=True + ) + return ranked + + def _generate_issue_recommendations( + self, + categorized_issues: Dict[str, List[Dict[str, Any]]], + severity_scores: Dict[str, Dict[str, Any]] + ) -> List[str]: + """Generate recommendations for addressing issues.""" + recommendations = [] + + for category, score_data in severity_scores.items(): + if score_data['priority'] == 'critical': + recommendations.append( + f"URGENT: Address {category} issues immediately - affecting {score_data['percentage']}% of reviews" + ) + elif score_data['priority'] == 'high': + recommendations.append( + f"HIGH PRIORITY: Focus on {category} issues in next update" + ) + + return recommendations + + def _extract_feature_request_text(self, text: str) -> str: + """Extract the specific feature request from review text.""" + # Simple extraction - find sentence with feature request keywords + sentences = text.split('.') + for sentence in sentences: + if any(keyword in sentence for keyword in self.FEATURE_REQUEST_KEYWORDS): + return sentence.strip() + return text[:100] # Fallback + + def _cluster_feature_requests( + self, + feature_requests: List[Dict[str, Any]] + ) -> List[Dict[str, Any]]: + """Cluster similar feature requests.""" + # Simplified clustering - group by common keywords + clusters = {} + + for request in feature_requests: + text = request['request_text'].lower() + # Extract key words + words = [w for w in text.split() if len(w) > 4] + + # Try to find matching cluster + matched = False + for cluster_key in clusters: + if any(word in cluster_key for word in words[:3]): + clusters[cluster_key].append(request) + matched = True + break + + if not matched and words: + cluster_key = ' '.join(words[:2]) + clusters[cluster_key] = [request] + + return [ + {'feature_theme': theme, 'request_count': len(requests), 'examples': requests[:3]} + for theme, requests in clusters.items() + ] + + def _prioritize_feature_requests( + self, + clustered_requests: List[Dict[str, Any]] + ) -> List[Dict[str, Any]]: + """Prioritize feature requests by frequency.""" + return sorted( + clustered_requests, + key=lambda x: x['request_count'], + reverse=True + )[:10] + + def _generate_feature_recommendations( + self, + prioritized_requests: List[Dict[str, Any]] + ) -> List[str]: + """Generate recommendations for feature requests.""" + recommendations = [] + + if prioritized_requests: + top_request = prioritized_requests[0] + recommendations.append( + f"Most requested feature: {top_request['feature_theme']} " + f"({top_request['request_count']} mentions) - consider for next major release" + ) + + if len(prioritized_requests) > 1: + recommendations.append( + f"Also consider: {prioritized_requests[1]['feature_theme']}" + ) + + return recommendations + + def _determine_trend_direction( + self, + rating_change: float, + sentiment_change: float + ) -> str: + """Determine overall trend direction.""" + if rating_change > 0.2 and sentiment_change > 5: + return 'improving' + elif rating_change < -0.2 and sentiment_change < -5: + return 'declining' + else: + return 'stable' + + def _generate_trend_insights( + self, + trends: List[Dict[str, Any]], + trend_direction: str + ) -> List[str]: + """Generate insights from trend analysis.""" + insights = [] + + if trend_direction == 'improving': + insights.append("Positive trend: User satisfaction is increasing over time") + elif trend_direction == 'declining': + insights.append("WARNING: User satisfaction is declining - immediate action needed") + else: + insights.append("Sentiment is stable - maintain current quality") + + # Review velocity insight + if len(trends) >= 2: + recent_reviews = trends[-1]['total_reviews'] + previous_reviews = trends[-2]['total_reviews'] + + if recent_reviews > previous_reviews * 1.5: + insights.append("Review volume increasing - growing user base or recent controversy") + + return insights + + +def analyze_reviews( + app_name: str, + reviews: List[Dict[str, Any]] +) -> Dict[str, Any]: + """ + Convenience function to perform comprehensive review analysis. + + Args: + app_name: App name + reviews: List of review dictionaries + + Returns: + Complete review analysis + """ + analyzer = ReviewAnalyzer(app_name) + + return { + 'sentiment_analysis': analyzer.analyze_sentiment(reviews), + 'common_themes': analyzer.extract_common_themes(reviews), + 'issues_identified': analyzer.identify_issues(reviews), + 'feature_requests': analyzer.find_feature_requests(reviews) + } diff --git a/skills/atlassian-admin/SKILL.md b/skills/atlassian-admin/SKILL.md new file mode 100644 index 00000000..5f05f285 --- /dev/null +++ b/skills/atlassian-admin/SKILL.md @@ -0,0 +1,234 @@ +--- +name: "atlassian-admin" +description: Atlassian Administrator for managing and organizing Atlassian products (Jira, Confluence, Bitbucket, Trello), users, permissions, security, integrations, system configuration, and org-wide governance. Use when asked to add users to Jira, change Confluence permissions, configure access control, update admin settings, manage Atlassian groups, set up SSO, install marketplace apps, review security policies, or handle any org-wide Atlassian administration task. +--- + +# Atlassian Administrator Expert + +## Workflows + +### User Provisioning +1. Create user account: `admin.atlassian.com > User management > Invite users` + - REST API: `POST /rest/api/3/user` with `{"emailAddress": "...", "displayName": "...","products": [...]}` +2. Add to appropriate groups: `admin.atlassian.com > User management > Groups > [group] > Add members` +3. Assign product access (Jira, Confluence) via `admin.atlassian.com > Products > [product] > Access` +4. Configure default permissions per group scheme +5. Send welcome email with onboarding info +6. **NOTIFY**: Relevant team leads of new member +7. **VERIFY**: Confirm user appears active at `admin.atlassian.com/o/{orgId}/users` and can log in + +### User Deprovisioning +1. **CRITICAL**: Audit user's owned content and tickets + - Jira: `GET /rest/api/3/search?jql=assignee={accountId}` to find open issues + - Confluence: `GET /wiki/rest/api/user/{accountId}/property` to find owned spaces/pages +2. Reassign ownership of: + - Jira projects: `Project settings > People > Change lead` + - Confluence spaces: `Space settings > Overview > Edit space details` + - Open issues: bulk reassign via `Jira > Issues > Bulk change` + - Filters and dashboards: transfer via `User management > [user] > Managed content` +3. Remove from all groups: `admin.atlassian.com > User management > [user] > Groups` +4. Revoke product access +5. Deactivate account: `admin.atlassian.com > User management > [user] > Deactivate` + - REST API: `DELETE /rest/api/3/user?accountId={accountId}` +6. **VERIFY**: Confirm `GET /rest/api/3/user?accountId={accountId}` returns `"active": false` +7. Document deprovisioning in audit log +8. **USE**: Jira Expert to reassign any remaining issues + +### Group Management +1. Create groups: `admin.atlassian.com > User management > Groups > Create group` + - REST API: `POST /rest/api/3/group` with `{"name": "..."}` + - Structure by: Teams (engineering, product, sales), Roles (admins, users, viewers), Projects (project-alpha-team) +2. Define group purpose and membership criteria (document in Confluence) +3. Assign default permissions per group +4. Add users to appropriate groups +5. **VERIFY**: Confirm group members via `GET /rest/api/3/group/member?groupName={name}` +6. Regular review and cleanup (quarterly) +7. **USE**: Confluence Expert to document group structure + +### Permission Scheme Design +**Jira Permission Schemes** (`Jira Settings > Issues > Permission Schemes`): +- **Public Project**: All users can view, members can edit +- **Team Project**: Team members full access, stakeholders view +- **Restricted Project**: Named individuals only +- **Admin Project**: Admins only + +**Confluence Permission Schemes** (`Confluence Admin > Space permissions`): +- **Public Space**: All users view, space members edit +- **Team Space**: Team-specific access +- **Personal Space**: Individual user only +- **Restricted Space**: Named individuals and groups + +**Best Practices**: +- Use groups, not individual permissions +- Principle of least privilege +- Regular permission audits +- Document permission rationale + +### SSO Configuration +1. Choose identity provider (Okta, Azure AD, Google) +2. Configure SAML settings: `admin.atlassian.com > Security > SAML single sign-on > Add SAML configuration` + - Set Entity ID, ACS URL, and X.509 certificate from IdP +3. Test SSO with admin account (keep password login active during test) +4. Test with regular user account +5. Enable SSO for organization +6. Enforce SSO: `admin.atlassian.com > Security > Authentication policies > Enforce SSO` +7. Configure SCIM for auto-provisioning: `admin.atlassian.com > User provisioning > [IdP] > Enable SCIM` +8. **VERIFY**: Confirm SSO flow succeeds and audit logs show `saml.login.success` events +9. Monitor SSO logs: `admin.atlassian.com > Security > Audit log > filter: SSO` + +### Marketplace App Management +1. Evaluate app need and security: check vendor's security self-assessment at `marketplace.atlassian.com` +2. Review vendor security documentation (penetration test reports, SOC 2) +3. Test app in sandbox environment +4. Purchase or request trial: `admin.atlassian.com > Billing > Manage subscriptions` +5. Install app: `admin.atlassian.com > Products > [product] > Apps > Find new apps` +6. Configure app settings per vendor documentation +7. Train users on app usage +8. **VERIFY**: Confirm app appears in `GET /rest/plugins/1.0/` and health check passes +9. Monitor app performance and usage; review annually for continued need + +### System Performance Optimization +**Jira** (`Jira Settings > System`): +- Archive old projects: `Project settings > Archive project` +- Reindex: `Jira Settings > System > Indexing > Full re-index` +- Clean up unused workflows and schemes: `Jira Settings > Issues > Workflows` +- Monitor queue/thread counts: `Jira Settings > System > System info` + +**Confluence** (`Confluence Admin > Configuration`): +- Archive inactive spaces: `Space tools > Overview > Archive space` +- Remove orphaned pages: `Confluence Admin > Orphaned pages` +- Monitor index and cache: `Confluence Admin > Cache management` + +**Monitoring Cadence**: +- Daily health checks: `admin.atlassian.com > Products > [product] > Health` +- Weekly performance reports +- Monthly capacity planning +- Quarterly optimization reviews + +### Integration Setup +**Common Integrations**: +- **Slack**: `Jira Settings > Apps > Slack integration` — notifications for Jira and Confluence +- **GitHub/Bitbucket**: `Jira Settings > Apps > DVCS accounts` — link commits to issues +- **Microsoft Teams**: `admin.atlassian.com > Apps > Microsoft Teams` +- **Zoom**: Available via Marketplace app `zoom-for-jira` +- **Salesforce**: Via Marketplace app `salesforce-connector` + +**Configuration Steps**: +1. Review integration requirements and OAuth scopes needed +2. Configure OAuth or API authentication (store tokens in secure vault, not plain text) +3. Map fields and data flows +4. Test integration thoroughly with sample data +5. Document configuration in Confluence runbook +6. Train users on integration features +7. **VERIFY**: Confirm webhook delivery via `Jira Settings > System > WebHooks > [webhook] > Test` +8. Monitor integration health via app-specific dashboards + +## Global Configuration + +### Jira Global Settings (`Jira Settings > Issues`) +**Issue Types**: Create and manage org-wide issue types; define issue type schemes; standardize across projects +**Workflows**: Create global workflow templates via `Workflows > Add workflow`; manage workflow schemes +**Custom Fields**: Create org-wide custom fields at `Custom fields > Add custom field`; manage field configurations and context +**Notification Schemes**: Configure default notification rules; create custom notification schemes; manage email templates + +### Confluence Global Settings (`Confluence Admin`) +**Blueprints & Templates**: Create org-wide templates at `Configuration > Global Templates and Blueprints`; manage blueprint availability +**Themes & Appearance**: Configure org branding at `Configuration > Themes`; customize logos and colors +**Macros**: Enable/disable macros at `Configuration > Macro usage`; configure macro permissions + +### Security Settings (`admin.atlassian.com > Security`) +**Authentication**: +- Password policies: `Security > Authentication policies > Edit` +- Session timeout: `Security > Session duration` +- API token management: `Security > API token controls` + +**Data Residency**: Configure data location at `admin.atlassian.com > Data residency > Pin products` + +**Audit Logs**: `admin.atlassian.com > Security > Audit log` +- Enable comprehensive logging; export via `GET /admin/v1/orgs/{orgId}/audit-log` +- Retain per policy (minimum 7 years for SOC 2/GDPR compliance) + +## Governance & Policies + +### Access Governance +- Quarterly review of all user access: `admin.atlassian.com > User management > Export users` +- Verify user roles and permissions; remove inactive users +- Limit org admins to 2–3 individuals; audit admin actions monthly +- Require MFA for all admins: `Security > Authentication policies > Require 2FA` + +### Naming Conventions +**Jira**: Project keys 3–4 uppercase letters (PROJ, WEB); issue types Title Case; custom fields prefixed (CF: Story Points) +**Confluence**: Spaces use Team/Project prefix (TEAM: Engineering); pages descriptive and consistent; labels lowercase, hyphen-separated + +### Change Management +**Major Changes**: Announce 2 weeks in advance; test in sandbox; create rollback plan; execute during off-peak; post-implementation review +**Minor Changes**: Announce 48 hours in advance; document in change log; monitor for issues + +## Disaster Recovery + +### Backup Strategy +**Jira & Confluence**: Daily automated backups; weekly manual verification; 30-day retention; offsite storage +- Trigger manual backup: `Jira Settings > System > Backup system` / `Confluence Admin > Backup and Restore` + +**Recovery Testing**: Quarterly recovery drills; document procedures; measure RTO and RPO + +### Incident Response +**Severity Levels**: +- **P1 (Critical)**: System down — respond in 15 min +- **P2 (High)**: Major feature broken — respond in 1 hour +- **P3 (Medium)**: Minor issue — respond in 4 hours +- **P4 (Low)**: Enhancement — respond in 24 hours + +**Response Steps**: +1. Acknowledge and log incident +2. Assess impact and severity +3. Communicate status to stakeholders +4. Investigate root cause (check `admin.atlassian.com > Products > [product] > Health` and Atlassian Status Page) +5. Implement fix +6. **VERIFY**: Confirm resolution via affected user test and health check +7. Post-mortem and lessons learned + +## Metrics & Reporting + +**System Health**: Active users (daily/weekly/monthly), storage utilization, API rate limits, integration health, response times +- Export via: `GET /admin/v1/orgs/{orgId}/users` for user counts; product-specific analytics dashboards + +**Usage Analytics**: Most active projects/spaces, content creation trends, user engagement, search patterns +**Compliance Metrics**: User access review completion, security audit findings, failed login attempts, API token usage + +## Decision Framework & Handoff Protocols + +**Escalate to Atlassian Support**: System outage, performance degradation org-wide, data loss/corruption, license/billing issues, complex migrations + +**Delegate to Product Experts**: +- Jira Expert: Project-specific configuration +- Confluence Expert: Space-specific settings +- Scrum Master: Team workflow needs +- Senior PM: Strategic planning input + +**Involve Security Team**: Security incidents, unusual access patterns, compliance audit preparation, new integration security review + +**TO Jira Expert**: New global workflows, custom fields, permission schemes, or automation capabilities available +**TO Confluence Expert**: New global templates, space permission schemes, blueprints, or macros configured +**TO Senior PM**: Usage analytics, capacity planning insights, cost optimization, security compliance status +**TO Scrum Master**: Team access provisioned, board configuration options, automation rules, integrations enabled +**FROM All Roles**: User access requests, permission changes, app installation requests, configuration support, incident reports + +## Atlassian MCP Integration + +**Primary Tools**: Jira MCP, Confluence MCP + +**Admin Operations**: +- User and group management via API +- Bulk permission updates +- Configuration audits +- Usage reporting +- System health monitoring +- Automated compliance checks + +**Integration Points**: +- Support all roles with admin capabilities +- Enable Jira Expert with global configurations +- Provide Confluence Expert with template management +- Ensure Senior PM has visibility into org health +- Enable Scrum Master with team provisioning diff --git a/skills/atlassian-admin/_meta.json b/skills/atlassian-admin/_meta.json new file mode 100644 index 00000000..7a882a29 --- /dev/null +++ b/skills/atlassian-admin/_meta.json @@ -0,0 +1,17 @@ +{ + "owner": "alirezarezvani", + "slug": "atlassian-admin", + "displayName": "atlassian-admin", + "latest": { + "version": "1.0.0", + "publishedAt": 1773243790654, + "commit": "https://github.com/openclaw/skills/commit/d6b7e4ec17803a4719360ef1299dc7788a4460ef" + }, + "history": [ + { + "version": "2.1.1", + "publishedAt": 1773193903487, + "commit": "https://github.com/openclaw/skills/commit/bf20838a561966bdf00ad2cd52838c3cb01467f8" + } + ] +} diff --git a/skills/atlassian-admin/assets/permission_scheme_template.json b/skills/atlassian-admin/assets/permission_scheme_template.json new file mode 100644 index 00000000..77ef7d78 --- /dev/null +++ b/skills/atlassian-admin/assets/permission_scheme_template.json @@ -0,0 +1,173 @@ +{ + "permissionScheme": { + "name": "Standard Project Permission Scheme", + "description": "Default permission scheme for standard projects. Assigns permissions based on project roles.", + "version": "1.0", + "lastUpdated": "YYYY-MM-DD", + "owner": "IT Admin Team" + }, + "roles": { + "projectAdmin": { + "description": "Full project administration including configuration and user management", + "typicalGroups": ["project-leads", "engineering-managers"] + }, + "developer": { + "description": "Create and manage issues, transitions, and attachments", + "typicalGroups": ["dept-engineering", "dept-product"] + }, + "user": { + "description": "View issues, add comments, and create basic issues", + "typicalGroups": ["org-all-employees"] + }, + "viewer": { + "description": "Read-only access to project issues and boards", + "typicalGroups": ["stakeholders", "external-contractors"] + } + }, + "permissions": { + "project": { + "ADMINISTER_PROJECTS": { + "description": "Manage project settings, roles, and permissions", + "grantedTo": ["projectAdmin"] + }, + "BROWSE_PROJECTS": { + "description": "View the project and its issues", + "grantedTo": ["projectAdmin", "developer", "user", "viewer"] + }, + "VIEW_DEV_TOOLS": { + "description": "View development panel (commits, branches, PRs)", + "grantedTo": ["projectAdmin", "developer"] + }, + "VIEW_READONLY_WORKFLOW": { + "description": "View read-only workflow", + "grantedTo": ["projectAdmin", "developer", "user", "viewer"] + } + }, + "issues": { + "CREATE_ISSUES": { + "description": "Create new issues in the project", + "grantedTo": ["projectAdmin", "developer", "user"] + }, + "EDIT_ISSUES": { + "description": "Edit issue fields", + "grantedTo": ["projectAdmin", "developer"] + }, + "DELETE_ISSUES": { + "description": "Delete issues permanently", + "grantedTo": ["projectAdmin"] + }, + "ASSIGN_ISSUES": { + "description": "Assign issues to team members", + "grantedTo": ["projectAdmin", "developer"] + }, + "ASSIGNABLE_USER": { + "description": "Be assigned to issues", + "grantedTo": ["projectAdmin", "developer"] + }, + "CLOSE_ISSUES": { + "description": "Close/resolve issues", + "grantedTo": ["projectAdmin", "developer"] + }, + "RESOLVE_ISSUES": { + "description": "Set issue resolution", + "grantedTo": ["projectAdmin", "developer"] + }, + "TRANSITION_ISSUES": { + "description": "Transition issues through workflow", + "grantedTo": ["projectAdmin", "developer", "user"] + }, + "LINK_ISSUES": { + "description": "Create and remove issue links", + "grantedTo": ["projectAdmin", "developer"] + }, + "MOVE_ISSUES": { + "description": "Move issues between projects", + "grantedTo": ["projectAdmin"] + }, + "SCHEDULE_ISSUES": { + "description": "Set due dates on issues", + "grantedTo": ["projectAdmin", "developer"] + }, + "SET_ISSUE_SECURITY": { + "description": "Set security level on issues", + "grantedTo": ["projectAdmin"] + } + }, + "comments": { + "ADD_COMMENTS": { + "description": "Add comments to issues", + "grantedTo": ["projectAdmin", "developer", "user"] + }, + "EDIT_ALL_COMMENTS": { + "description": "Edit any comment", + "grantedTo": ["projectAdmin"] + }, + "EDIT_OWN_COMMENTS": { + "description": "Edit own comments", + "grantedTo": ["projectAdmin", "developer", "user"] + }, + "DELETE_ALL_COMMENTS": { + "description": "Delete any comment", + "grantedTo": ["projectAdmin"] + }, + "DELETE_OWN_COMMENTS": { + "description": "Delete own comments", + "grantedTo": ["projectAdmin", "developer", "user"] + } + }, + "attachments": { + "CREATE_ATTACHMENTS": { + "description": "Attach files to issues", + "grantedTo": ["projectAdmin", "developer", "user"] + }, + "DELETE_ALL_ATTACHMENTS": { + "description": "Delete any attachment", + "grantedTo": ["projectAdmin"] + }, + "DELETE_OWN_ATTACHMENTS": { + "description": "Delete own attachments", + "grantedTo": ["projectAdmin", "developer", "user"] + } + }, + "worklogs": { + "WORK_ON_ISSUES": { + "description": "Log work on issues", + "grantedTo": ["projectAdmin", "developer"] + }, + "EDIT_ALL_WORKLOGS": { + "description": "Edit any worklog", + "grantedTo": ["projectAdmin"] + }, + "EDIT_OWN_WORKLOGS": { + "description": "Edit own worklogs", + "grantedTo": ["projectAdmin", "developer"] + }, + "DELETE_ALL_WORKLOGS": { + "description": "Delete any worklog", + "grantedTo": ["projectAdmin"] + }, + "DELETE_OWN_WORKLOGS": { + "description": "Delete own worklogs", + "grantedTo": ["projectAdmin", "developer"] + } + } + }, + "projectMappings": [ + { + "projectKey": "EXAMPLE", + "projectName": "Example Project", + "scheme": "Standard Project Permission Scheme", + "roleAssignments": { + "projectAdmin": ["project-leads"], + "developer": ["team-example-devs"], + "user": ["org-all-employees"], + "viewer": ["stakeholders-example"] + } + } + ], + "notes": { + "usage": "Copy this template and customize role assignments per project. Use group names that match your Atlassian groups.", + "review": "Review permission scheme assignments quarterly as part of access review.", + "changes": "Any changes to permission schemes should be documented and approved by IT Admin." + } +} diff --git a/skills/atlassian-admin/references/security-hardening-guide.md b/skills/atlassian-admin/references/security-hardening-guide.md new file mode 100644 index 00000000..66ea0864 --- /dev/null +++ b/skills/atlassian-admin/references/security-hardening-guide.md @@ -0,0 +1,214 @@ +# Atlassian Cloud Security Hardening Guide + +## Overview + +This guide provides a comprehensive security hardening checklist for Atlassian Cloud products (Jira, Confluence, Bitbucket). It covers identity management, access controls, data protection, and monitoring practices aligned with enterprise security standards. + +## Identity & Authentication + +### SSO / SAML Setup + +**Implementation Steps:** +1. Verify your domain in Atlassian Admin (admin.atlassian.com) +2. Claim all company email accounts +3. Configure SAML SSO with your identity provider (Okta, Azure AD, Google Workspace) +4. Set authentication policy to enforce SSO for all managed accounts +5. Test with a pilot group before full rollout +6. Disable password-based login for managed accounts + +**Configuration Checklist:** +- [ ] Domain verified and accounts claimed +- [ ] SAML IdP configured with correct entity ID and SSO URL +- [ ] Attribute mapping: email, displayName, groups +- [ ] Single Logout (SLO) configured +- [ ] Authentication policy enforcing SSO +- [ ] Fallback access configured for emergency admin accounts +- [ ] SCIM provisioning enabled for automatic user sync + +### Two-Factor Authentication (2FA) + +**Enforcement Policy:** +- [ ] 2FA required for all managed accounts +- [ ] Enforce via authentication policy (not just recommended) +- [ ] Hardware security keys (FIDO2/WebAuthn) preferred for admin accounts +- [ ] TOTP (authenticator app) as minimum for all users +- [ ] SMS-based 2FA disabled (SIM swap vulnerability) +- [ ] Recovery codes generated and stored securely + +### Session Management +- [ ] Session timeout set to 8 hours of inactivity (maximum) +- [ ] Absolute session timeout: 24 hours +- [ ] Require re-authentication for sensitive operations +- [ ] Monitor concurrent sessions per user +- [ ] Enforce session termination on password change + +## Access Controls + +### IP Allowlisting + +**Configuration:** +- [ ] Enable IP allowlisting for organization +- [ ] Add corporate office IP ranges +- [ ] Add VPN exit node IP addresses +- [ ] Add CI/CD server IPs for API access +- [ ] Test access from all approved locations +- [ ] Document approved IP ranges with justification +- [ ] Review IP allowlist quarterly + +**Exceptions:** +- Mobile access may require VPN or MDM solution +- Remote workers need VPN or conditional access policies +- API integrations need stable IP ranges + +### API Token Management + +**Policies:** +- [ ] Inventory all API tokens in use +- [ ] Set maximum token lifetime (90 days recommended) +- [ ] Require token rotation on schedule +- [ ] Use service accounts for integrations (not personal tokens) +- [ ] Monitor API token usage patterns +- [ ] Revoke tokens immediately on employee departure +- [ ] Document purpose and owner for each token + +**Best Practices:** +- Use OAuth 2.0 (3LO) for user-context integrations +- Use API tokens only for service-to-service +- Store tokens in secrets management (never in code) +- Implement least-privilege scopes for OAuth apps + +### Permission Model +- [ ] Review global permissions quarterly +- [ ] Use groups for permission assignment (not individual users) +- [ ] Implement role-based access for Jira projects +- [ ] Restrict Confluence space admin to designated owners +- [ ] Limit Jira system admin to 2-3 people +- [ ] Audit "anyone" or "logged in users" permissions +- [ ] Remove direct user permissions where groups exist + +## Audit & Monitoring + +### Audit Log Configuration + +**What to Monitor:** +- User authentication events (login, logout, failed attempts) +- Permission changes (project, space, global) +- User account changes (creation, deactivation, group changes) +- API token creation and revocation +- App installations and updates +- Data export operations +- Admin configuration changes + +**Setup Steps:** +- [ ] Enable organization audit log +- [ ] Configure audit log retention (minimum 1 year) +- [ ] Set up automated export to SIEM (Splunk, Datadog, etc.) +- [ ] Create alerts for suspicious patterns +- [ ] Schedule monthly audit log review +- [ ] Document incident response procedures for alerts + +### Alerting Rules + +**Critical Alerts (Immediate Response):** +- Multiple failed login attempts (>5 in 10 minutes) +- Admin permission grants to unexpected users +- API token created by non-service accounts +- Bulk data export or deletion +- New third-party app installed with broad permissions + +**Warning Alerts (Same-Day Review):** +- New admin users added +- Permission scheme changes +- Authentication policy modifications +- IP allowlist changes +- User deactivation (verify it is expected) + +## Data Protection + +### Data Residency +- [ ] Configure data residency realm (US, EU, AU, etc.) +- [ ] Verify product data pinned to selected region +- [ ] Document data residency for compliance audits +- [ ] Review data residency coverage (some metadata may be global) +- [ ] Monitor for new residency options from Atlassian + +### Encryption +- [ ] Verify encryption at rest (AES-256, managed by Atlassian) +- [ ] Verify encryption in transit (TLS 1.2+) +- [ ] Review Atlassian's encryption key management practices +- [ ] Consider BYOK (Bring Your Own Key) for Atlassian Guard Premium + +### Data Loss Prevention +- [ ] Configure content restrictions for sensitive pages/issues +- [ ] Implement classification labels (public, internal, confidential) +- [ ] Restrict file attachment types if needed +- [ ] Monitor bulk exports and downloads +- [ ] Set up DLP rules for sensitive data patterns (PII, credentials) + +## Mobile Device Management + +### Mobile Access Controls +- [ ] Require MDM enrollment for mobile Atlassian apps +- [ ] Enforce device encryption +- [ ] Require screen lock with biometrics or PIN +- [ ] Enable remote wipe capability +- [ ] Block rooted/jailbroken devices +- [ ] Restrict copy/paste to managed apps +- [ ] Set app-level PIN for Atlassian apps + +### Mobile Policies +- [ ] Define approved mobile devices/OS versions +- [ ] Enforce automatic app updates +- [ ] Configure offline data access limits +- [ ] Set maximum offline cache duration +- [ ] Review mobile access logs monthly + +## Third-Party App Security + +### App Review Process +- [ ] Maintain approved app list (whitelist) +- [ ] Review app permissions before installation +- [ ] Verify app is Atlassian Marketplace certified +- [ ] Check app vendor security certifications +- [ ] Assess data access scope (read-only vs read-write) +- [ ] Review app privacy policy +- [ ] Document app owner and business justification + +### App Governance +- [ ] Audit installed apps quarterly +- [ ] Remove unused apps (no usage in 90 days) +- [ ] Monitor app permission changes +- [ ] Restrict app installation to admins only +- [ ] Review Atlassian Guard app access policies +- [ ] Set up alerts for new app installations + +## Compliance Documentation + +### Required Documentation +- [ ] Security policy for Atlassian Cloud usage +- [ ] Access control matrix (roles, permissions, justification) +- [ ] Incident response plan for Atlassian security events +- [ ] Data classification policy applied to Atlassian content +- [ ] Third-party app risk assessments +- [ ] Annual security review report + +### Compliance Frameworks +- **SOC 2:** Map Atlassian controls to Trust Service Criteria +- **ISO 27001:** Align with Annex A controls for cloud services +- **GDPR:** Configure data residency, right to deletion, DPAs +- **HIPAA:** Review BAA availability, encryption, access controls + +## Hardening Schedule + +| Task | Frequency | Owner | +|------|-----------|-------| +| Permission audit | Quarterly | IT Admin | +| API token rotation | Every 90 days | Integration owners | +| App review | Quarterly | IT Admin | +| Audit log review | Monthly | Security team | +| IP allowlist review | Quarterly | IT Admin | +| Authentication policy review | Semi-annually | Security team | +| Full security assessment | Annually | Security team | +| User access review | Quarterly | Managers + IT Admin | +| Data residency verification | Annually | Compliance | +| Mobile device audit | Quarterly | IT Admin | diff --git a/skills/atlassian-admin/references/user-provisioning-checklist.md b/skills/atlassian-admin/references/user-provisioning-checklist.md new file mode 100644 index 00000000..d0bbf552 --- /dev/null +++ b/skills/atlassian-admin/references/user-provisioning-checklist.md @@ -0,0 +1,177 @@ +# User Provisioning & Lifecycle Management Checklist + +## Overview + +This checklist covers the complete user lifecycle in Atlassian Cloud products, from onboarding through offboarding. Consistent provisioning ensures security, compliance, and a smooth user experience. + +## Onboarding Steps + +### Pre-Provisioning +- [ ] Receive approved access request (ticket or HR system trigger) +- [ ] Verify employee record in HR system +- [ ] Determine role-based access level (see Role Templates below) +- [ ] Identify required Atlassian products (Jira, Confluence, Bitbucket) +- [ ] Identify required project/space access + +### Account Creation +- [ ] User account auto-provisioned via SCIM (preferred) or manually created +- [ ] Email domain matches verified organization domain +- [ ] SSO authentication verified (user can log in via IdP) +- [ ] 2FA enrollment confirmed +- [ ] Correct product access assigned (Jira, Confluence, Bitbucket) + +### Group Membership +- [ ] Add to organization-level groups (e.g., `all-employees`) +- [ ] Add to department group (e.g., `engineering`, `product`, `marketing`) +- [ ] Add to team-specific groups (e.g., `team-platform`, `team-mobile`) +- [ ] Add to project groups as needed (e.g., `project-alpha-members`) +- [ ] Verify group membership grants correct permissions + +### Product Configuration +- [ ] **Jira:** Add to correct project roles (Developer, User, Admin) +- [ ] **Jira:** Assign to correct board(s) +- [ ] **Jira:** Set default dashboard if applicable +- [ ] **Confluence:** Grant access to relevant spaces +- [ ] **Confluence:** Add to space groups with appropriate permission level +- [ ] **Bitbucket:** Grant repository access per team +- [ ] **Bitbucket:** Configure branch permissions + +### Welcome & Training +- [ ] Send welcome email with access details and key links +- [ ] Share Confluence onboarding page (getting started guide) +- [ ] Assign onboarding buddy for Atlassian tool questions +- [ ] Schedule optional training session for new users +- [ ] Provide link to internal Atlassian usage guidelines + +## Role-Based Access Templates + +### Developer +- **Jira:** Project Developer role (create, edit, transition issues) +- **Confluence:** Team space editor, documentation spaces viewer +- **Bitbucket:** Repository write access for team repos + +### Product Manager +- **Jira:** Project Admin role (manage boards, workflows, components) +- **Confluence:** Product spaces editor, all team spaces viewer +- **Bitbucket:** Repository read access (optional) + +### Designer +- **Jira:** Project User role (view, comment, transition) +- **Confluence:** Design space editor, product spaces editor +- **Bitbucket:** No access (unless needed) + +### Engineering Manager +- **Jira:** Project Admin for managed projects, viewer for others +- **Confluence:** Team space admin, all spaces viewer +- **Bitbucket:** Repository admin for team repos + +### Executive / Stakeholder +- **Jira:** Viewer role on strategic projects, dashboard access +- **Confluence:** Viewer on relevant spaces +- **Bitbucket:** No access + +### Contractor / External +- **Jira:** Project User role, limited to specific projects +- **Confluence:** Viewer on specific spaces only (no edit) +- **Bitbucket:** Repository read access, specific repos only +- **Additional:** Set account expiration date, restrict IP access + +## Group Membership Standards + +### Naming Convention +``` +org-{company} # Organization-wide groups +dept-{department} # Department groups +team-{team-name} # Team-specific groups +project-{project} # Project-scoped groups +role-{role} # Role-based groups (role-admin, role-viewer) +``` + +### Standard Groups +| Group | Purpose | Products | +|-------|---------|----------| +| `org-all-employees` | All full-time employees | Jira, Confluence | +| `dept-engineering` | All engineers | Jira, Confluence, Bitbucket | +| `dept-product` | All product team | Jira, Confluence | +| `dept-marketing` | All marketing team | Confluence | +| `role-jira-admins` | Jira administrators | Jira | +| `role-confluence-admins` | Confluence administrators | Confluence | +| `role-org-admins` | Organization administrators | All | + +## Offboarding Procedure + +### Immediate Actions (Day of Departure) +- [ ] Deactivate user account in Atlassian (or via IdP/SCIM) +- [ ] Revoke all API tokens associated with the user +- [ ] Revoke all OAuth app authorizations +- [ ] Transfer ownership of critical Confluence pages +- [ ] Reassign Jira issues (open/in-progress items) +- [ ] Remove from all groups +- [ ] Document access removal in offboarding ticket + +### Within 24 Hours +- [ ] Verify account is fully deactivated (cannot log in) +- [ ] Check for shared credentials or service accounts +- [ ] Review audit log for recent activity +- [ ] Transfer Confluence space ownership if applicable +- [ ] Update Jira project leads/component leads if applicable +- [ ] Remove from any Atlassian Marketplace vendor accounts + +### Within 7 Days +- [ ] Verify no lingering sessions or cached access +- [ ] Review integrations the user may have set up +- [ ] Check for automation rules owned by the user +- [ ] Update team dashboards and filters +- [ ] Confirm with manager that all transfers are complete + +### Data Retention +- [ ] User content (pages, issues, comments) retained per policy +- [ ] Personal spaces archived or transferred +- [ ] Account marked as deactivated (not deleted) for audit trail +- [ ] Data deletion request processed if required (GDPR) + +## Quarterly Access Reviews + +### Review Process +1. Generate user access report from Atlassian Admin +2. Distribute to managers for team verification +3. Managers confirm or flag each user's access level +4. IT Admin processes approved changes +5. Document review completion for compliance + +### Review Checklist +- [ ] All active accounts match current employee list +- [ ] No accounts for departed employees +- [ ] Group memberships align with current roles +- [ ] Admin access limited to approved administrators +- [ ] External/contractor accounts have valid expiration dates +- [ ] Service accounts documented with current owners +- [ ] Unused accounts (no login in 90 days) flagged for review + +### Compliance Documentation +- [ ] Access review completion date recorded +- [ ] Manager sign-off captured (email or ticket) +- [ ] Changes made during review documented +- [ ] Exceptions documented with justification and approval +- [ ] Report filed for audit purposes +- [ ] Next review date scheduled + +## Automation Opportunities + +### SCIM Provisioning +- Automatically create/deactivate accounts based on IdP changes +- Sync group membership from IdP groups +- Reduce manual provisioning errors +- Ensure immediate deactivation on termination + +### Workflow Automation +- Trigger onboarding checklist from HR system event +- Auto-assign to groups based on department/role attributes +- Send welcome messages via Confluence automation +- Schedule access reviews via Jira recurring tickets + +### Monitoring +- Alert on accounts without 2FA after 7 days +- Alert on admin group changes +- Weekly report of new and deactivated accounts +- Monthly stale account report (no login in 90 days) diff --git a/skills/atlassian-admin/scripts/permission_audit_tool.py b/skills/atlassian-admin/scripts/permission_audit_tool.py new file mode 100644 index 00000000..ee24c480 --- /dev/null +++ b/skills/atlassian-admin/scripts/permission_audit_tool.py @@ -0,0 +1,469 @@ +#!/usr/bin/env python3 +""" +Permission Audit Tool + +Analyzes Atlassian permission schemes for security issues. Checks for +over-permissioned groups, direct user permissions, missing restrictions on +sensitive actions, inconsistencies across projects, and compliance gaps. + +Usage: + python permission_audit_tool.py permissions.json + python permission_audit_tool.py permissions.json --format json +""" + +import argparse +import json +import sys +from typing import Any, Dict, List, Optional, Set + + +# --------------------------------------------------------------------------- +# Audit Configuration +# --------------------------------------------------------------------------- + +SENSITIVE_PERMISSIONS = { + "administer_project", + "administer_jira", + "delete_issues", + "delete_all_comments", + "delete_all_attachments", + "manage_watchers", + "modify_reporter", + "bulk_change", + "system_admin", + "manage_group_filter_subscriptions", +} + +RECOMMENDED_GROUP_ONLY_PERMISSIONS = { + "browse_projects", + "create_issues", + "edit_issues", + "transition_issues", + "assign_issues", + "resolve_issues", + "close_issues", + "add_comments", + "edit_all_comments", +} + +SEVERITY_WEIGHTS = { + "critical": 25, + "high": 15, + "medium": 8, + "low": 3, + "info": 1, +} + + +# --------------------------------------------------------------------------- +# Audit Checks +# --------------------------------------------------------------------------- + +def check_over_permissioned_groups( + schemes: List[Dict[str, Any]], +) -> List[Dict[str, str]]: + """Check for groups with overly broad admin access.""" + findings = [] + + for scheme in schemes: + scheme_name = scheme.get("name", "Unknown Scheme") + grants = scheme.get("grants", []) + + group_permissions = {} + for grant in grants: + group = grant.get("group", "") + permission = grant.get("permission", "").lower() + if group: + if group not in group_permissions: + group_permissions[group] = set() + group_permissions[group].add(permission) + + for group, perms in group_permissions.items(): + admin_perms = perms & SENSITIVE_PERMISSIONS + if len(admin_perms) >= 3: + findings.append({ + "rule": "over_permissioned_group", + "severity": "high", + "scheme": scheme_name, + "group": group, + "message": f"Group '{group}' has {len(admin_perms)} sensitive permissions " + f"in scheme '{scheme_name}': {', '.join(sorted(admin_perms))}. " + f"Review if all are necessary.", + }) + + if "system_admin" in perms or "administer_jira" in perms: + findings.append({ + "rule": "admin_access_warning", + "severity": "critical", + "scheme": scheme_name, + "group": group, + "message": f"Group '{group}' has system/Jira admin access in '{scheme_name}'. " + f"Ensure this is strictly necessary and membership is limited.", + }) + + return findings + + +def check_direct_user_permissions( + schemes: List[Dict[str, Any]], +) -> List[Dict[str, str]]: + """Check for permissions granted directly to users instead of groups.""" + findings = [] + + for scheme in schemes: + scheme_name = scheme.get("name", "Unknown Scheme") + grants = scheme.get("grants", []) + + for grant in grants: + user = grant.get("user", "") + permission = grant.get("permission", "") + + if user and not grant.get("group"): + severity = "high" if permission.lower() in SENSITIVE_PERMISSIONS else "medium" + findings.append({ + "rule": "direct_user_permission", + "severity": severity, + "scheme": scheme_name, + "user": user, + "message": f"User '{user}' has direct permission '{permission}' in '{scheme_name}'. " + f"Use groups instead for maintainability and audit clarity.", + }) + + return findings + + +def check_missing_restrictions( + schemes: List[Dict[str, Any]], +) -> List[Dict[str, str]]: + """Check for missing restrictions on sensitive actions.""" + findings = [] + + for scheme in schemes: + scheme_name = scheme.get("name", "Unknown Scheme") + grants = scheme.get("grants", []) + + granted_permissions = set() + for grant in grants: + granted_permissions.add(grant.get("permission", "").lower()) + + # Check if delete permissions are unrestricted + delete_perms = {"delete_issues", "delete_all_comments", "delete_all_attachments"} + unrestricted_deletes = delete_perms & granted_permissions + + for grant in grants: + perm = grant.get("permission", "").lower() + group = grant.get("group", "") + if perm in delete_perms and group: + # Check if granted to broad groups + broad_groups = {"users", "everyone", "all-users", "jira-users", "jira-software-users"} + if group.lower() in broad_groups: + findings.append({ + "rule": "unrestricted_delete", + "severity": "critical", + "scheme": scheme_name, + "message": f"Delete permission '{perm}' granted to broad group '{group}' " + f"in '{scheme_name}'. Restrict to admins or leads only.", + }) + + # Check if admin permissions exist + admin_perms = {"administer_project", "administer_jira", "system_admin"} + if not (admin_perms & granted_permissions): + findings.append({ + "rule": "no_admin_defined", + "severity": "medium", + "scheme": scheme_name, + "message": f"No explicit admin permission defined in '{scheme_name}'. " + f"Ensure project administration is properly assigned.", + }) + + return findings + + +def check_scheme_consistency( + schemes: List[Dict[str, Any]], +) -> List[Dict[str, str]]: + """Check for inconsistencies across permission schemes.""" + findings = [] + + if len(schemes) < 2: + return findings + + # Compare permission sets across schemes + scheme_perms = {} + for scheme in schemes: + name = scheme.get("name", "Unknown") + perms = set() + for grant in scheme.get("grants", []): + perms.add(grant.get("permission", "").lower()) + scheme_perms[name] = perms + + # Find schemes with significantly different permission sets + all_perms = set() + for perms in scheme_perms.values(): + all_perms |= perms + + scheme_names = list(scheme_perms.keys()) + for i in range(len(scheme_names)): + for j in range(i + 1, len(scheme_names)): + name_a = scheme_names[i] + name_b = scheme_names[j] + diff = scheme_perms[name_a].symmetric_difference(scheme_perms[name_b]) + if len(diff) > 5: + findings.append({ + "rule": "scheme_inconsistency", + "severity": "medium", + "message": f"Schemes '{name_a}' and '{name_b}' differ significantly " + f"({len(diff)} different permissions). Review for intentional differences.", + }) + + return findings + + +def check_compliance_gaps( + schemes: List[Dict[str, Any]], +) -> List[Dict[str, str]]: + """Check for common compliance gaps.""" + findings = [] + + for scheme in schemes: + scheme_name = scheme.get("name", "Unknown Scheme") + grants = scheme.get("grants", []) + + groups_used = set() + users_used = set() + for grant in grants: + if grant.get("group"): + groups_used.add(grant["group"]) + if grant.get("user"): + users_used.add(grant["user"]) + + # Check for separation of duties + admin_groups = set() + for grant in grants: + if grant.get("permission", "").lower() in SENSITIVE_PERMISSIONS and grant.get("group"): + admin_groups.add(grant["group"]) + + if len(admin_groups) == 1 and len(groups_used) > 1: + findings.append({ + "rule": "separation_of_duties", + "severity": "info", + "scheme": scheme_name, + "message": f"Only one group ('{next(iter(admin_groups))}') holds all sensitive permissions " + f"in '{scheme_name}'. Consider separating duties across multiple groups.", + }) + + # Check user count + if len(users_used) > 5: + findings.append({ + "rule": "too_many_direct_users", + "severity": "high", + "scheme": scheme_name, + "message": f"Scheme '{scheme_name}' has {len(users_used)} direct user grants. " + f"Migrate to group-based permissions for better governance.", + }) + + return findings + + +# --------------------------------------------------------------------------- +# Main Analysis +# --------------------------------------------------------------------------- + +def audit_permissions(data: Dict[str, Any]) -> Dict[str, Any]: + """Run full permission audit.""" + schemes = data.get("schemes", []) + + if not schemes: + # Try treating the entire input as a single scheme + if data.get("grants") or data.get("name"): + schemes = [data] + else: + return { + "risk_score": 0, + "grade": "invalid", + "error": "No permission schemes found in input", + "findings": [], + "summary": {}, + } + + all_findings = [] + all_findings.extend(check_over_permissioned_groups(schemes)) + all_findings.extend(check_direct_user_permissions(schemes)) + all_findings.extend(check_missing_restrictions(schemes)) + all_findings.extend(check_scheme_consistency(schemes)) + all_findings.extend(check_compliance_gaps(schemes)) + + # Calculate risk score (higher = more risk) + summary = {"critical": 0, "high": 0, "medium": 0, "low": 0, "info": 0} + total_penalty = 0 + for finding in all_findings: + severity = finding["severity"] + summary[severity] = summary.get(severity, 0) + 1 + total_penalty += SEVERITY_WEIGHTS.get(severity, 0) + + risk_score = min(100, total_penalty) + health_score = max(0, 100 - risk_score) + + if health_score >= 85: + grade = "excellent" + elif health_score >= 70: + grade = "good" + elif health_score >= 50: + grade = "fair" + else: + grade = "poor" + + # Generate remediation recommendations + remediations = _generate_remediations(all_findings) + + return { + "risk_score": risk_score, + "health_score": health_score, + "grade": grade, + "schemes_analyzed": len(schemes), + "findings": all_findings, + "summary": summary, + "remediations": remediations, + } + + +def _generate_remediations(findings: List[Dict[str, str]]) -> List[str]: + """Generate remediation recommendations.""" + remediations = [] + rules_seen = set() + + for finding in findings: + rule = finding["rule"] + if rule in rules_seen: + continue + rules_seen.add(rule) + + if rule == "over_permissioned_group": + remediations.append("Review and reduce sensitive permissions for over-permissioned groups. Apply principle of least privilege.") + elif rule == "admin_access_warning": + remediations.append("Audit admin group membership. Limit system/Jira admin access to essential personnel only.") + elif rule == "direct_user_permission": + remediations.append("Migrate direct user permissions to group-based grants. Create functional groups for common permission sets.") + elif rule == "unrestricted_delete": + remediations.append("Restrict delete permissions to project admins or leads. Remove from broad user groups.") + elif rule == "scheme_inconsistency": + remediations.append("Standardize permission schemes across projects. Document intentional differences.") + elif rule == "too_many_direct_users": + remediations.append("Create groups for users with direct permissions. This simplifies onboarding/offboarding.") + elif rule == "separation_of_duties": + remediations.append("Consider splitting admin responsibilities across multiple groups for better separation of duties.") + elif rule == "no_admin_defined": + remediations.append("Define explicit admin permissions in each scheme to ensure proper project governance.") + + return remediations + + +# --------------------------------------------------------------------------- +# Output Formatting +# --------------------------------------------------------------------------- + +def format_text_output(result: Dict[str, Any]) -> str: + """Format results as readable text report.""" + lines = [] + lines.append("=" * 60) + lines.append("PERMISSION AUDIT REPORT") + lines.append("=" * 60) + lines.append("") + + if "error" in result: + lines.append(f"ERROR: {result['error']}") + return "\n".join(lines) + + lines.append("AUDIT SUMMARY") + lines.append("-" * 30) + lines.append(f"Risk Score: {result['risk_score']}/100 (lower is better)") + lines.append(f"Health Score: {result['health_score']}/100") + lines.append(f"Grade: {result['grade'].title()}") + lines.append(f"Schemes Analyzed: {result['schemes_analyzed']}") + lines.append("") + + summary = result.get("summary", {}) + lines.append("FINDINGS BY SEVERITY") + lines.append("-" * 30) + lines.append(f"Critical: {summary.get('critical', 0)}") + lines.append(f"High: {summary.get('high', 0)}") + lines.append(f"Medium: {summary.get('medium', 0)}") + lines.append(f"Low: {summary.get('low', 0)}") + lines.append(f"Info: {summary.get('info', 0)}") + lines.append("") + + findings = result.get("findings", []) + if findings: + lines.append("DETAILED FINDINGS") + lines.append("-" * 30) + for i, finding in enumerate(findings, 1): + severity = finding["severity"].upper() + lines.append(f"{i}. [{severity}] {finding['message']}") + lines.append(f" Rule: {finding['rule']}") + if finding.get("scheme"): + lines.append(f" Scheme: {finding['scheme']}") + lines.append("") + + remediations = result.get("remediations", []) + if remediations: + lines.append("REMEDIATION RECOMMENDATIONS") + lines.append("-" * 30) + for i, rem in enumerate(remediations, 1): + lines.append(f"{i}. {rem}") + + return "\n".join(lines) + + +def format_json_output(result: Dict[str, Any]) -> Dict[str, Any]: + """Format results as JSON.""" + return result + + +# --------------------------------------------------------------------------- +# CLI Interface +# --------------------------------------------------------------------------- + +def main() -> int: + """Main CLI entry point.""" + parser = argparse.ArgumentParser( + description="Audit Atlassian permission schemes for security issues" + ) + parser.add_argument( + "permissions_file", + help="JSON file with permission scheme data", + ) + parser.add_argument( + "--format", + choices=["text", "json"], + default="text", + help="Output format (default: text)", + ) + + args = parser.parse_args() + + try: + with open(args.permissions_file, "r") as f: + data = json.load(f) + + result = audit_permissions(data) + + if args.format == "json": + print(json.dumps(format_json_output(result), indent=2)) + else: + print(format_text_output(result)) + + return 0 + + except FileNotFoundError: + print(f"Error: File '{args.permissions_file}' not found", file=sys.stderr) + return 1 + except json.JSONDecodeError as e: + print(f"Error: Invalid JSON in '{args.permissions_file}': {e}", file=sys.stderr) + return 1 + except Exception as e: + print(f"Error: {e}", file=sys.stderr) + return 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/skills/atlassian-templates/SKILL.md b/skills/atlassian-templates/SKILL.md new file mode 100644 index 00000000..7f8455b7 --- /dev/null +++ b/skills/atlassian-templates/SKILL.md @@ -0,0 +1,252 @@ +--- +name: "atlassian-templates" +description: Atlassian Template and Files Creator/Modifier expert for creating, modifying, and managing Jira and Confluence templates, blueprints, custom layouts, reusable components, and standardized content structures. Use when building org-wide templates, custom blueprints, page layouts, and automated content generation. +--- + +# Atlassian Template & Files Creator Expert + +Specialist in creating, modifying, and managing reusable templates and files for Jira and Confluence. Ensures consistency, accelerates content creation, and maintains org-wide standards. + +--- + +## Workflows + +### Template Creation Process +1. **Discover**: Interview stakeholders to understand needs +2. **Analyze**: Review existing content patterns +3. **Design**: Create template structure and placeholders +4. **Implement**: Build template with macros and formatting +5. **Test**: Validate with sample data — confirm template renders correctly in preview before publishing +6. **Document**: Create usage instructions +7. **Publish**: Deploy to appropriate space/project via MCP (see MCP Operations below) +8. **Verify**: Confirm deployment success; roll back to previous version if errors occur +9. **Train**: Educate users on template usage +10. **Monitor**: Track adoption and gather feedback +11. **Iterate**: Refine based on usage + +### Template Modification Process +1. **Assess**: Review change request and impact +2. **Version**: Create new version, keep old available +3. **Modify**: Update template structure/content +4. **Test**: Validate changes don't break existing usage; preview updated template before publishing +5. **Migrate**: Provide migration path for existing content +6. **Communicate**: Announce changes to users +7. **Support**: Assist users with migration +8. **Archive**: Deprecate old version after transition; confirm deprecated template is unlisted, not deleted + +### Blueprint Development +1. Define blueprint scope and purpose +2. Design multi-page structure +3. Create page templates for each section +4. Configure page creation rules +5. Add dynamic content (Jira queries, user data) +6. Test blueprint creation flow end-to-end with a sample space +7. Verify all macro references resolve correctly before deployment +8. **HANDOFF TO**: Atlassian Admin for global deployment + +--- + +## Confluence Templates Library + +See **TEMPLATES.md** for full reference tables and copy-paste-ready template structures. The following summarises the standard types this skill creates and maintains. + +### Confluence Template Types +| Template | Purpose | Key Macros Used | +|----------|---------|-----------------| +| **Meeting Notes** | Structured meeting records with agenda, decisions, and action items | `{date}`, `{tasks}`, `{panel}`, `{info}`, `{note}` | +| **Project Charter** | Org-level project scope, stakeholder RACI, timeline, and budget | `{panel}`, `{status}`, `{timeline}`, `{info}` | +| **Sprint Retrospective** | Agile ceremony template with What Went Well / Didn't Go Well / Actions | `{panel}`, `{expand}`, `{tasks}`, `{status}` | +| **PRD** | Feature definition with goals, user stories, functional/non-functional requirements, and release plan | `{panel}`, `{status}`, `{jira}`, `{warning}` | +| **Decision Log** | Structured option analysis with decision matrix and implementation tracking | `{panel}`, `{status}`, `{info}`, `{tasks}` | + +**Standard Sections** included across all Confluence templates: +- Header panel with metadata (owner, date, status) +- Clearly labelled content sections with inline placeholder instructions +- Action items block using `{tasks}` macro +- Related links and references + +### Complete Example: Meeting Notes Template + +The following is a copy-paste-ready Meeting Notes template in Confluence storage format (wiki markup): + +``` +{panel:title=Meeting Metadata|borderColor=#0052CC|titleBGColor=#0052CC|titleColor=#FFFFFF} +*Date:* {date} +*Owner / Facilitator:* @[facilitator name] +*Attendees:* @[name], @[name] +*Status:* {status:colour=Yellow|title=In Progress} +{panel} + +h2. Agenda +# [Agenda item 1] +# [Agenda item 2] +# [Agenda item 3] + +h2. Discussion & Decisions +{panel:title=Key Decisions|borderColor=#36B37E|titleBGColor=#36B37E|titleColor=#FFFFFF} +* *Decision 1:* [What was decided and why] +* *Decision 2:* [What was decided and why] +{panel} + +{info:title=Notes} +[Detailed discussion notes, context, or background here] +{info} + +h2. Action Items +{tasks} +* [ ] [Action item] — Owner: @[name] — Due: {date} +* [ ] [Action item] — Owner: @[name] — Due: {date} +{tasks} + +h2. Next Steps & Related Links +* Next meeting: {date} +* Related pages: [link] +* Related Jira issues: {jira:key=PROJ-123} +``` + +> Full examples for all other template types (Project Charter, Sprint Retrospective, PRD, Decision Log) and all Jira templates can be generated on request or found in **TEMPLATES.md**. + +--- + +## Jira Templates Library + +### Jira Template Types +| Template | Purpose | Key Sections | +|----------|---------|--------------| +| **User Story** | Feature requests in As a / I want / So that format | Acceptance Criteria (Given/When/Then), Design links, Technical Notes, Definition of Done | +| **Bug Report** | Defect capture with reproduction steps | Environment, Steps to Reproduce, Expected vs Actual Behavior, Severity, Workaround | +| **Epic** | High-level initiative scope | Vision, Goals, Success Metrics, Story Breakdown, Dependencies, Timeline | + +**Standard Sections** included across all Jira templates: +- Clear summary line +- Acceptance or success criteria as checkboxes +- Related issues and dependencies block +- Definition of Done (for stories) + +--- + +## Macro Usage Guidelines + +**Dynamic Content**: Use macros for auto-updating content (dates, user mentions, Jira queries) +**Visual Hierarchy**: Use `{panel}`, `{info}`, and `{note}` to create visual distinction +**Interactivity**: Use `{expand}` for collapsible sections in long templates +**Integration**: Embed Jira charts and tables via `{jira}` macro for live data + +--- + +## Atlassian MCP Integration + +**Primary Tools**: Confluence MCP, Jira MCP + +### Template Operations via MCP + +All MCP calls below use the exact parameter names expected by the Atlassian MCP server. Replace angle-bracket placeholders with real values before executing. + +**Create a Confluence page template:** +```json +{ + "tool": "confluence_create_page", + "parameters": { + "space_key": "PROJ", + "title": "Template: Meeting Notes", + "body": "", + "labels": ["template", "meeting-notes"], + "parent_id": "" + } +} +``` + +**Update an existing template:** +```json +{ + "tool": "confluence_update_page", + "parameters": { + "page_id": "", + "version": "", + "title": "Template: Meeting Notes", + "body": "", + "version_comment": "v2 — added status macro to header" + } +} +``` + +**Create a Jira issue description template (via field configuration):** +```json +{ + "tool": "jira_update_field_configuration", + "parameters": { + "project_key": "PROJ", + "field_id": "description", + "default_value": "