Technology

CLI tools vs MCP servers: a revolution in AI agent efficiency

Nx removed most of its MCP tools. Google is betting on the CLI. Why? Token efficiency, security and the SKILLS.md pattern. See tests showing up to 90% fewer tokens.

Mateusz Mieczysław Ciszczoń
Mateusz Mieczysław Ciszczoń Technology
A terminal with code showing CLI tools in action

Introduction: the MCP paradox

In 2024 Anthropic introduced the Model Context Protocol (MCP) as a standard for communication between AI agents and external tools. The promise was tempting: one protocol, thousands of integrations, seamless data exchange.

A year later? Nx – one of the largest users of MCP – removed most of its MCP tools. Google released Gemini CLI instead of full MCP integration. The developer community is talking louder and louder about "MCP overhead".

What happened?

Thesis: In most cases, ordinary CLI tools are more efficient, safer and cheaper in token terms than MCP servers. And this is not an opinion – there is hard data to back it up.

What is MCP and why was it created?

Model Context Protocol is an open standard created by Anthropic for communication between AI agents (e.g. Claude, ChatGPT) and external systems.

MCP architecture

┌─────────────┐         ┌─────────────┐         ┌──────────────┐
│   AI Agent  │◄────────┤  MCP Client │◄────────┤  MCP Server  │
│  (Claude)   │  JSON   │             │  stdio  │  (GitHub)    │
└─────────────┘         └─────────────┘         └──────────────┘

MCP consists of three elements:

  1. MCP server – a process responsible for communication with an external API (e.g. GitHub, Slack, file system)
  2. MCP client – an intermediary layer in an AI application (Claude Desktop, Cursor, Windsurf)
  3. Protocol – standard JSON-RPC exchange over stdio/HTTP/SSE

The MCP ecosystem (2025)

State as of March 2025:

  • 8,941+ repositories on GitHub tagged "mcp"
  • Support: Claude Desktop, ChatGPT, Cursor, Zed, Windsurf, VSCode (Copilot)
  • Official servers: GitHub, Google Drive, Slack, Postgres, Puppeteer, Brave Search
  • Community: over 300 unofficial integrations

Sounds impressive, right? Let's see where the problem lies.

Problem #1: token consumption

The biggest headache: MCP servers often return gigantic JSON payloads that gobble up the AI agent's context window.

Example: GitHub pull requests

Query: "Show me the names of open PRs in repository X"

MCP variant

// MCP server response (~500 tokens)
{
  "pulls": [
    {
      "id": 1234567890,
      "node_id": "PR_kwDOABCD123",
      "number": 42,
      "state": "open",
      "locked": false,
      "title": "Fix authentication bug",
      "user": {
        "login": "devuser",
        "id": 98765,
        "node_id": "MDQ6VXNlcjk4NzY1",
        "avatar_url": "https://...",
        "gravatar_id": "",
        "url": "https://api.github.com/users/devuser",
        "html_url": "https://github.com/devuser",
        // ... 20+ additional fields
      },
      "body": "This PR fixes a critical authentication bug...",
      "created_at": "2025-03-01T10:00:00Z",
      "updated_at": "2025-03-07T14:30:00Z",
      "closed_at": null,
      "merged_at": null,
      "merge_commit_sha": null,
      "assignee": { /* ... another 20+ fields ... */ },
      "assignees": [ /* ... */ ],
      "requested_reviewers": [ /* ... */ ],
      "labels": [ /* ... */ ],
      "milestone": null,
      "_links": { /* ... */ },
      // ... and so on
    },
    // ... more PRs
  ]
}

The agent gets: Full objects with dozens of unused fields. The context window fills up in an instant.

CLI variant

# CLI command (50 tokens of output)
gh pr list --state open --json number,title

[
  {"number": 42, "title": "Fix authentication bug"},
  {"number": 43, "title": "Add dark mode"},
  {"number": 44, "title": "Update dependencies"}
]

The agent gets: Only what it needs. 90% fewer tokens.

Real-world impact: the Nx monorepo

The Nx team tested its MCP tools on various tasks. The results:

  • Simple question and answer: MCP marginally more expensive (more JSON overhead)
  • Complex tasks: MCP significantly more expensive (large data payloads × many calls)
  • Token trade-off: The Skills pattern consumes more tokens at the start (the agent reads documentation) but delivers better results (accuracy, use of generators, verification)

Key conclusion: It is not just about the raw number of tokens – it is about the quality of task execution in proportion to the tokens consumed.

Problem #2: security and deployment complexity

MCP = a running server

Each MCP server is a separate process running permanently in the background:

# A typical MCP server
node mcp-github-server.js
  ↳ Listening on port 3000...
  ↳ Websocket server started...
  ↳ Auth middleware active...

Security implications:

  1. Open port (stdio/HTTP/SSE) – an additional attack surface
  2. Persistent process – must be supervised (systemd, pm2, Docker)
  3. Credential management – tokens/keys must be available to a long-running process
  4. Network isolation – communication between servers must be managed

CLI = process isolation

CLI tools work on a "run and exit" model:

# CLI command
gh api /repos/owner/repo/pulls --jq '.[].title'
  ↳ Process starts
  ↳ Executes an HTTP request
  ↳ Returns JSON
  ↳ Process exits (total time: 0.3s)

Security advantages:

  1. No persistent process – there is nothing left to attack after it finishes
  2. Operating-system-level permissions – the OS manages access
  3. Credential management – tokens in the system keychain (macOS Keychain, Linux Secret Service)
  4. Sandboxing – easy to run in an isolated environment (container, VM)

Deployment comparison

AspectMCP serverCLI tools
Installationnpm install + server configurationbrew/apt install (ready-made)
AuthenticationCustom per server (env variables, config files)Standard OS keychain
MonitoringLogs, health checks, restart policiesNot needed
UpdatesManually per server, breaking protocol changesPackage manager (brew upgrade)
Resource usageMemory/CPU per server (3-20+ servers = overhead)Zero when unused

Quote from Nx: "Why would we maintain 15 MCP servers when gh, git, npm are already installed and working?"

CLI tools: the natural alternative

Why is the CLI natural for AI agents?

Modern AI agents (Claude Sonnet 4.5, GPT-5, Gemini Pro) are natively trained to use the CLI. They don't need any additional "translator" – they simply know how to:

# Executed by the agent (a real example from Claude)
$ git log --oneline --since="2 weeks ago" | wc -l
42

$ npm run test -- --changed

$ docker ps --filter "status=running" --format "{{.Names}}"

Large language models have seen billions of lines of CLI documentation, man pages and StackOverflow examples during training. This is their native language.

Ecosystem: what already exists

Most development environments already have installed:

ToolPurposeAI-friendly output
ghGitHub APIJSON (--json flag)
gitVersion controlParseable text
npm/yarn/pnpmPackage managementJSON (--json)
dockerContainersJSON (--format)
kubectlKubernetesJSON/YAML (-o json)
jqJSON processingFiltered JSON
curlHTTP requestsRaw response
psqlPostgreSQLCSV/JSON output

API stability: These tools have years of backward compatibility. A git from 2015 still works identically. MCP? Breaking changes every few months.

CLI + jq: precise extraction

The combination of CLI + jq allows for surgical data extraction:

# Only active CI/CD pipelines from the last 24h
gh api /repos/owner/repo/actions/runs \
  --jq '[.workflow_runs[]
    | select(.status == "in_progress")
    | select(.created_at > (now - 86400 | todate))
    | {id, name, status, created_at}]'

Output: Precisely what you need. Zero overhead. Zero extra tokens.

The SKILLS.md pattern: teaching instead of tools

Key discovery: AI agents perform better when given domain knowledge instead of raw tools.

What is SKILLS.md?

SKILLS.md is a documentation file for a specific tool created especially for AI agents. Instead of:

"Here's a function listPullRequests(repo, state) – call it"

You say:

"WHEN you want to see open PRs, use gh pr list --state open --json number,title. HOW to filter: --author, --label, --search. WHY: It returns minimal JSON – saves tokens. EXAMPLES: [3 real use cases]"

Skills vs MCP: the Anthropic approach

Think of it as the difference between:

  • MCP: Giving someone a key (a tool)
  • Skills: Teaching them how to be a mechanic (knowledge)

Data from the Nx tests:

ModelTask typeMCP accuracySkills accuracyWinner
Claude SonnetGraph navigation87%92%Skills +5%
Claude HaikuUsing generators71%89%Skills +18%
GPT-4CI fixing83%88%Skills +5%
Gemini ProMulti-step tasks79%85%Skills +6%

The biggest difference for smaller models (Haiku, GPT-5-mini) – the Skills pattern compensates for a lack of compute power through better education.

A real example from Nx: the generators skill

# SKILL: Nx Generators

## When to use

When the user wants to:
- Create a new library/component/application
- Generate a standard code structure
- Follow workspace conventions

## How to do it

```bash
# List available generators
nx list

# Show generator options
nx g @nx/react:component --help

# Run with a dry-run first
nx g @nx/react:component mycomp --dry-run

# Run for real
nx g @nx/react:component mycomp

Why it works

  • --dry-run prevents errors (the AI can verify before executing)
  • Generator names follow patterns (the AI can infer)
  • Parseable JSON output: --json | jq '.changes'

Common patterns

  1. Feature module: nx g @nx/react:lib feature-auth --directory=libs/features
  2. UI component: nx g @nx/react:component Button --project=shared-ui
  3. E2E test: nx g @nx/cypress:e2e my-app-e2e

Error handling

If a generator fails, check:

  • Does the target project exist? nx show project <name>
  • Is the generator installed? nx list @nx/react
**Result**: The agent knows **WHEN**, **HOW** and **WHY** to use the generator. Not just "that it can".

Incremental loading

Skills are loaded on demand:

# The agent does not see ALL skills at the start
# Only when it needs them:

Agent: "I need to work with GitHub PRs"
  ↓
System: [loads github-pr.skills.md]
  ↓
Agent: [reads 200 tokens of domain knowledge]
  ↓
Agent: "Ah, I should use `gh pr list --json`"

MCP: The agent gets all available functions at the start (listing resources = 1,000+ tokens).

Skills: The agent gets only what it needs (incrementally = 50-200 tokens per skill).

Case study: Nx.dev – "Why We Deleted Most of Our MCP Tools"

In January 2025, Nx published the article "Why We Deleted Most of Our MCP Tools" – the best case study of the MCP → Skills + CLI transition.

Context: full MCP adoption → removal

2024 Q3: Nx built a full MCP integration:

  • MCP servers for: workspace operations, project graph, generators, migrations, CI analytics
  • Integration in Claude Desktop + Cursor + Windsurf
  • The officially recommended way to work with Nx through AI

2025 Q1: Nx removed most of its MCP tools.

What remained: Only 2 MCP servers for authenticated APIs and running processes.

Why? 3 main reasons

1. A paradigm shift: "ask and answer" → "plan and execute"

The old era (2024): AI agents as a "smart chat" – you ask, you get an answer, you copy it into your IDE.

The new era (2025): AI agents as "autonomous developers" (Claude Code, Cursor Composer) – they plan, execute, commit.

The MCP problem: Designed for "ask and answer". Not for "autonomous execution".

The CLI solution: Agents can natively execute commands – that is their core skill.

2. The token economics did not make sense

Nx tested on real production codebases (Nx Cloud, large enterprise monorepos).

Test data:

Task: "List all projects affected by changes"

MCP approach:
  1. Call listProjects() → 2,400 tokens (all projects with metadata)
  2. Call getGraph() → 1,800 tokens (full dependency graph)
  3. Call listChanges() → 600 tokens (git diff metadata)
  Total: 4,800 tokens + agent processing

CLI approach:
  1. nx show projects --affected → 180 tokens (names only)
  2. (if needed) nx show project <name> --json → 120 tokens per project
  Total: 180-600 tokens depending on the depth needed

Savings: 87-96% fewer tokens

Quote from the article: "We were paying to transmit data we didn't need."

3. Skills outperformed MCP (empirically)

Nx conducted controlled experiments comparing MCP vs Skills:

Test setup:

  • 50 tasks per approach (generators, CI fixes, refactoring, testing)
  • 4 models: Claude Sonnet, Haiku, GPT-4, Gemini Pro
  • Metrics: accuracy (correct output), appropriate tool use, verification steps

Results (aggregated):

MetricMCPSkillsDelta
Task accuracy81%88%+7%
Generator use62%79%+17%
Verification performed43%71%+28%
Token cost (simple)450380-16%
Token cost (complex)1,2001,650+38%

Interpretation:

  • Accuracy: Skills win through domain knowledge (the agent understands WHY)
  • Generator use: Skills teach WHEN to use generators (MCP only teaches "that you can")
  • Verification: The Skills pattern encourages --dry-run, previewing, testing
  • Token cost: Simple tasks are cheaper with MCP (less overhead), complex ones are more expensive with Skills but produce better results

Key quote: "We don't optimise for tokens. We optimise for correct task execution. Skills win."

The biggest impact: smaller models

Claude Haiku (a fast, cheap model):

  • MCP accuracy: 71%
  • Skills accuracy: 89%
  • Delta: +18%

Why? Haiku has less compute power – it needs more guidance. The Skills pattern gives it that.

Sonnet/GPT-4 are also better with Skills, but the delta is smaller (+5-7%) – stronger models cope with raw tools.

Business implication: The Skills pattern allows the use of cheaper models with better results.

What remained in MCP?

Nx did not remove everything. Two use cases remained:

1. Authenticated APIs

// MCP: Nx Cloud API (private, authenticated)
{
  "command": "node",
  "args": ["nx-cloud-mcp-server.js"],
  "env": {
    "NX_CLOUD_TOKEN": "***"
  }
}

Why MCP? The Nx Cloud API requires:

  • An OAuth flow
  • Rate limiting
  • Webhook handling
  • Real-time updates (SSE)

CLI alternative? Possible, but it would require building your own CLI tool from scratch. MCP gives it "out of the box".

2. Running processes

// MCP: Development server monitoring
{
  "command": "node",
  "args": ["nx-dev-server-mcp.js"]
}

Use case: The agent monitors the nx serve output, restarts on error, coordinates hot reload.

Why MCP? Communication with a long-running process – the CLI exits after execution, MCP keeps the connection open.

The Nx rule: "Skills for knowledge, MCP for connectivity"

Final architecture:

┌─────────────────────────────────────────────┐
│         AI Agent (Claude/GPT)               │
└───────────┬────────────────────┬────────────┘
            │                    │
    ┌───────▼───────┐    ┌──────▼──────────┐
    │  Skills.md    │    │  MCP (minimal)  │
    │  (knowledge)  │    │  (connectivity) │
    └───────┬───────┘    └──────┬──────────┘
            │                    │
    ┌───────▼───────┐    ┌──────▼──────────┐
    │  CLI tools    │    │  Nx Cloud API   │
    │  nx, git, jq  │    │  Dev servers    │
    └───────────────┘    └─────────────────┘

Decision rule:

  1. Can it be done in the CLI? → Use Skills + CLI
  2. Is authentication/a long-running process needed? → Use MCP
  3. A borderline case? → Default to CLI (simpler)

When to use what: a decision framework

Based on the Nx case study + community data, here is a decision tree:

Use CLI + Skills when:

  1. The tool already exists (gh, docker, kubectl, npm, git)
  2. Data can be filtered (jq, grep, awk are enough)
  3. No authentication is needed OR standard auth (OS keychain)
  4. Short-lived operations (seconds to minutes)
  5. Simple parsing (JSON, CSV, plain text)

Examples:

  • GitHub operations → use the gh CLI
  • git operations → use git commands
  • Package management → use npm/yarn
  • Container operations → use the docker CLI
  • Cloud operations → use the aws/gcloud/az CLI

Use MCP when:

  1. Complex authentication (OAuth, SAML, multi-step)
  2. Real-time updates are needed (webhooks, SSE, websockets)
  3. A stateful connection is required (database pooling, sessions)
  4. Long-running processes (monitoring, dev servers)
  5. The CLI tool does not exist AND building one is complex

Examples:

  • Private authenticated APIs (Nx Cloud, internal tools)
  • Real-time monitoring (dev servers, CI/CD pipelines)
  • Database connections (connection pooling matters)
  • Complex API flows (Slack bots, multi-step forms)

Comparison matrix

DimensionCLI + SkillsMCP serverWinner
Token efficiency80-95% less (filtered output)Full payloads (JSON dumps)CLI
SecurityProcess isolation, OS permissionsRunning server, open portCLI
DeploymentAlready installedInstallation + config per serverCLI
AuthenticationOS keychain (standard)Custom per serverCLI
Real-time updatesPolling requiredNative SSE/websocketsMCP
Stateful operationsRun and exitPersistent connectionMCP
API stabilityYears of backward compatibilityFrequent breaking changesCLI
Learning curve (AI)Native (trained on CLI)Requires learning the MCP protocolCLI
CommunityDecades of tools/documentationGrowing (8,941+ repositories)CLI
Nx empirical data88% accuracy, more generators81% accuracyCLI/Skills

Result: CLI 9 – MCP 2

Not "CLI vs MCP" – rather "CLI + MCP".

Gemini CLI example:

# Google's hybrid approach
gemini-cli configure
  ↓
Uses: CLI for local operations (reading files, git)
       MCP for the Gemini API (authenticated, real-time)
       Skills for domain knowledge (when to use what)

Nx architecture:

// nx-workspace.skills.md teaches agents:
// "Use the `nx` CLI for 99% of operations"
// "Use MCP only for the authenticated Nx Cloud API"

Result: The best of both worlds – efficiency where possible, connectivity where needed.

Implementation guide: adding CLI support for AI agents

Do you have an existing CLI tool and want to make it "AI-agent-friendly"? Here's how.

1. A JSON output flag

Add a --json flag for every command:

// Before (human-readable)
$ my-tool list-items
✓ Item 1: Active
✓ Item 2: Pending
✓ Item 3: Completed

// After (machine-readable)
$ my-tool list-items --json
[
  {"id": 1, "name": "Item 1", "status": "active"},
  {"id": 2, "name": "Item 2", "status": "pending"},
  {"id": 3, "name": "Item 3", "status": "completed"}
]

Implementation (Node.js example):

import { Command } from "commander";

const programme = new Command();

programme
  .command("list-items")
  .option("--json", "Output as JSON")
  .action(async (options) => {
    const items = await fetchItems();

    if (options.json) {
      console.log(JSON.stringify(items, null, 2));
    } else {
      items.forEach((item) => {
        console.log(`✓ ${item.name}: ${item.status}`);
      });
    }
  });

Why it matters: AI agents can natively parse JSON – no external dependencies needed.

2. Minimal output by default

Problem: CLI tools often dump everything.

# Bad: Returns everything (3,000 tokens)
$ gh api /repos/owner/repo/pulls

# Good: Returns only what is needed (200 tokens)
$ gh pr list --json number,title,author

Rule: The default output should be minimal – the user can add --verbose if they need more.

3. Composability with jq/grep/awk

Design the CLI output so it works with pipelines:

# Composability
$ my-tool list --json | jq '.[] | select(.status=="active")'
$ my-tool list --json | grep "important"
$ my-tool stats | awk '{sum += $2} END {print sum}'

Anti-pattern: A custom query language inside the CLI – unnecessary, use jq.

4. Error handling for agents

A standard error format:

// Error output (JSON) - always to stderr
{
  "error": "Authentication failed",
  "code": "AUTH_ERROR",
  "hint": "Run `my-tool login` to authenticate"
}

// Success output (JSON) - always to stdout
[{ "id": 1, "name": "Item 1" }]

Why separate streams? The agent can:

# Capture the output
result=$(my-tool list --json 2>/dev/null)

# Capture errors
errors=$(my-tool list --json 2>&1 >/dev/null)

5. Create SKILLS.md

File structure:

# SKILL: [Tool name]

## When to use

[Describe use cases - WHEN the agent should use this tool]

## Commands

### [command-name]

**Purpose**: [What it does]
**Usage**: `my-tool [command] [flags]`
**Output**: [Example JSON structure]

**Flags**:
- `--json`: JSON output
- `--filter`: Filter results

**Examples**:
[3-5 real examples]

## Common patterns

[Frequent combinations, pipelines, workflows]

## Troubleshooting

[Common errors + how to fix them]

Where to put it: docs/SKILLS.md in the CLI tool's repository – agents will be able to read it via GitHub.

6. Security: sandboxed execution

AI agents execute arbitrary commands – safeguards are needed:

// Wrapper: Safe CLI execution
import { exec } from "child_process";
import { promisify } from "util";

const execAsync = promisify(exec);

async function executeCliSafely(command: string) {
  // Whitelist of allowed commands
  const allowedCommands = ["gh", "git", "npm", "my-tool"];
  const cmd = command.split(" ")[0];

  if (!allowedCommands.includes(cmd)) {
    throw new Error(`Command not allowed: ${cmd}`);
  }

  // Timeout protection
  const timeout = 30000; // 30s

  // Execute in an isolated env
  const result = await execAsync(command, {
    timeout,
    env: {
      // Pass only the necessary env variables
      PATH: process.env.PATH,
      HOME: process.env.HOME,
    },
    maxBuffer: 1024 * 1024, // max 1MB of output
  });

  return result.stdout;
}

Security checklist:

  • Whitelist of allowed commands
  • Timeout limits (preventing hangs)
  • Output size limits (preventing memory exhaustion)
  • Environment isolation (minimal env variables)
  • No shell expansion (use exec, not eval)

7. Testing AI-friendliness

Test scenarios:

// Test 1: Parseable JSON output
const output = await executeCliSafely("my-tool list --json");
const parsed = JSON.parse(output); // Should not throw
assert(Array.isArray(parsed));

// Test 2: Composability
const filtered = await executeCliSafely(
  "my-tool list --json | jq '.[] | select(.status==\"active\")'"
);
assert(filtered.length > 0);

// Test 3: Error handling
try {
  await executeCliSafely("my-tool list --invalid-flag");
} catch (err) {
  const error = JSON.parse(err.stderr);
  assert(error.code === "INVALID_FLAG");
}

CI integration: Run these tests in GitHub Actions – every PR must pass them.

The future: convergence, not competition

Thesis: CLI and MCP do not compete – they converge into a hybrid architecture.

Trend #1: MCP is becoming lighter

Anthropic's recommendation (December 2024):

"Instead of returning full objects, consider returning minimal data and letting agents run CLI commands for details."

Example: The GitHub MCP server new version (v2.0):

// Old (v1.0): Return everything
{
  "tool": "github_list_prs",
  "returns": [/* 50+ fields per PR */]
}

// New (v2.0): Return pointers + CLI
{
  "tool": "github_list_prs",
  "returns": [
    {"number": 42, "cli_command": "gh pr view 42 --json"}
  ]
}

Result: MCP as a catalogue + pointers, CLI as the execution layer.

Trend #2: CLI tools are adding MCP endpoints

The reverse move: CLI tools are adding optional MCP servers for advanced use cases.

Example: AWS CLI v3 (2025 roadmap):

# Traditional CLI
$ aws s3 ls

# New: MCP server mode
$ aws mcp-server start --port 3000
  ↳ Exposure: Real-time S3 events (PutObject, DeleteObject)
  ↳ Maintaining: CloudWatch log streaming
  ↳ Ensuring: Authenticated session (no re-auth per request)

Use case: Long-running operations where CLI execute-and-exit is not enough.

Trend #3: SKILLS.md as a standard

A growing movement: SKILLS.md as a documentation standard for AI agents.

Repositories already adopting it (March 2025):

  • Nx (nx-workspace.skills.md)
  • GitHub CLI (gh-cli.skills.md)
  • Docker CLI (docker.skills.md maintained by the community)
  • Kubernetes (kubectl.skills.md in k8s/website)

Proposal: A W3C standard "Agent Skills Specification" (draft stage) – a standard format for SKILLS.md.

Trend #4: IDE integration

VSCode Copilot (2025 insider builds):

  • Automatic discovery of SKILLS.md (in the workspace root)
  • CLI execution in the integrated terminal (Copilot can execute)
  • MCP fallback for APIs without a CLI

Cursor / Windsurf: A similar architecture – CLI first, MCP for authenticated services.

Result: The developer experience does not require manual configuration – everything is auto-detected.

Prediction: the state of play in 2026

By the end of 2026:

  1. 80% of AI agent operations via CLI + Skills
  2. 20% via MCP (authenticated APIs, real time)
  3. SKILLS.md widely adopted (a de-facto standard)
  4. The MCP protocol stabilised (v2.0 with breaking changes completed)
  5. Hybrid by default in major IDEs (VSCode, Cursor, Zed)

The winner: Developers – less overhead, more control, better results.

Summary: the key takeaways

1. CLI + Skills > MCP in most cases

Empirical data (Nx):

  • 88% accuracy vs 81% (Skills)
  • 80-95% fewer tokens (CLI)
  • Better security (process isolation)
  • Simpler deployment (already installed)

2. MCP is not bad – it is overused

MCP makes sense for:

  • Authenticated APIs (OAuth, multi-step)
  • Real-time communication (SSE, websockets)
  • Long-running processes (dev servers, monitoring)

MCP does not make sense for:

  • GitHub operations → use the gh CLI
  • git operations → use git commands
  • Cloud operations → use the aws/gcloud/az CLI
  • Package management → use npm/yarn

3. The Skills pattern = a game changer

Why Skills win:

  • They teach WHEN to use a tool (MCP only teaches "that you can")
  • Incremental loading (only what is needed)
  • Better for smaller models (+18% accuracy for Haiku)
  • Domain knowledge > raw tools

4. The hybrid approach is the way

Not "vs" but "+":

Skills.md → CLI → Execution (99% of cases)
                ↓
              MCP → Authenticated APIs (1% of cases)

The Nx rule: "Skills for knowledge, MCP for connectivity"

5. Your action points

If you are building a CLI tool:

  1. Add a --json flag
  2. Minimal output by default
  3. Create SKILLS.md
  4. Test with AI agents

If you are using MCP:

  1. Check whether the CLI is not enough
  2. If it is – switch to CLI + Skills
  3. If not – stay with MCP but opt-in

If you are configuring AI agents:

  1. Prefer CLI over MCP (where possible)
  2. Use MCP only where necessary
  3. Create SKILLS.md for custom tools

Resources

Official documentation

  1. Model Context Protocol (MCP)

  2. Anthropic research

  3. GitHub CLI

Case studies and articles

  1. Nx.dev: "Why We Deleted Most of Our MCP Tools"

    • Full article
    • Must-read – the best case study of the MCP → CLI transition
    • Includes: comparative tests, accuracy charts, production data
  2. Google Gemini CLI

Community and tools

  1. The MCP ecosystem

  2. SKILLS.md examples

  3. CLI tools for AI agents

    • jq – JSON processor
    • fx – interactive JSON
    • glow – Markdown renderer

Further reading

  1. Token efficiency research

  2. Security best practices

  3. The future of AI agents

Do you need efficient CLI tools?

Contact us

We will help you optimise your workflow with AI agents and build efficient CLI tools for your organisation.

What do you do? *
Newsletter

By submitting your enquiry you agree to our privacy policy.

Free consultation

Let's talk about your idea!

We offer free online consultations where we are happy to talk about your needs and ideas and help you determine the best direction.

Our technologies