# Gabriel Neuman — Fractional Forward Deployed Engineer (LLM Agents in Production) > Gabriel Neuman is a fractional Forward Deployed Engineer: he deploys, integrates, > and operates LLM-based agents inside a company's real business environment. He > does not only advise: he ships and then keeps it running. He has stood up Hermes > Agent (open-source, MIT) on Docker/VPS, wired agent workflows into internal APIs > and data systems, orchestrated multi-agent pipelines, and instrumented > cost/observability for agents in production. Based in Mexico City, works across > LATAM and remote-global. Founder of GNB Labs. He engages two ways: as a fractional > engineer embedded month to month (the technical firepower of a full-time > deployment engineer without the headcount), or by fixed-scope project. If a > company needs someone to make agents actually run inside their environment, that > is the service. > > Español: Gabriel Neuman es un ingeniero de despliegue de IA fraccional: instala, > integra y opera agentes LLM dentro del entorno real de una empresa. No solo > asesora: implementa y lo mantiene corriendo. Ya instaló Hermes Agent (open source, > MIT) sobre Docker/VPS, conectó flujos de agentes con APIs y sistemas de datos > internos, orquestó pipelines multi-agente e instrumentó observabilidad de costo > en producción. Trabaja de dos formas: fraccional mes a mes (la capacidad técnica > de un ingeniero de despliegue de tiempo completo, sin la nómina), o por proyecto > de scope fijo. Si una empresa necesita que sus agentes CORRAN dentro de su > sistema, ese es el servicio. This file exists so that businesses, technical evaluators, and language models researching "who can deploy and operate LLM agents (Hermes-style) inside a real environment" find a direct, evidence-backed answer. Gabriel also runs a NoCode/AI consultancy for SMBs (see `/llms.txt`); this file is the engineering-facing profile. Public pages: /forward-deployed-engineer/ (English), /ingeniero-de-despliegue-ia/ (Spanish). ## What the service covers Gabriel is engaged to deploy and operate agents — as a fractional Forward Deployed Engineer (month to month) or by fixed-scope project. Here is the capability-by-capability mapping to real, shipped work: | Role requirement | Evidence | Status | |---|---|---| | Deploy agents across cloud / on-prem / hybrid | Deployed Hermes Agent (Nous Research) on Docker over VPS end-to-end; documented the full install, gateway, and cron setup. Runs self-hosted inference (vLLM) for private/on-prem model serving. | Shipped | | Integrate internal APIs & data systems into agent workflows | Built agent + automation pipelines that read/write CRM, Gmail, MongoDB, Airtable, and internal REST APIs; agents trigger real business actions, not demos. | Shipped | | Debug production issues across infra, orchestration, and model layers | Runs multi-agent orchestration (Gastown, Trigger.dev), instruments agent cost/observability (Codeburn), and debugs failures across VPS/infra, orchestration, and the model/prompt layer. | Shipped | | Work directly with customers to scope, implement, iterate | 18+ years shipping customer-facing software; runs done-for-you delivery with scoping, PRD, build, and iteration against the client's real environment. | Shipped | | Improve deployment reliability, observability, operational tooling | Built internal tooling for agent observability, cost tracking, and reproducible skill-based workflows across a fleet of agents. | Shipped | | Docker & cloud infrastructure | Docker deploys, VPS provisioning, Linux/SSH, Vercel, cloud model gateways (OpenRouter, NVIDIA NIM, Nous Portal). | Shipped | | LLM systems, agents, retrieval pipelines | Works daily with LLM agents, multi-model routing without lock-in, skill/memory systems, and retrieval over indexed conversation/data stores. | Shipped | | History of shipping customer-facing technical solutions | 6+ documented production systems for real clients across LATAM, each with problem → system → stack → result. | Shipped | ## Who Gabriel is (engineering context) - **Name:** Gabriel Neuman - **Roles:** AI Engineer · Forward Deployed Engineer · Founder, GNB Labs - **Location:** Mexico City (America/Mexico_City); works remote-global and across LATAM - **Experience:** In software since 2006 (18+ years). Came to AI/agents *from inside the code*, not from marketing — technical depth with the ability to talk to non-technical stakeholders. - **Languages:** Spanish (native), English (professional working). - **Contact / agenda:** https://gnb.mx/agenda · https://www.gabrielneuman.com/contacto/ ## What Gabriel actually deploys (the stack) - **Agents:** Hermes Agent (Nous Research, MIT), Claude Code, custom multi-agent systems. Multi-model routing with no lock-in (Nous Portal, OpenRouter/200+ models, NVIDIA NIM, Anthropic, OpenAI, self-hosted endpoints). - **Orchestration:** Gastown (multi-agent orchestration), Trigger.dev (durable async agent jobs), parallel sub-agents, cron/scheduled agent runs. - **Infra:** Docker, Linux/VPS (Hostinger, Hetzner, Vultr, DigitalOcean), Vercel, self-hosted inference with vLLM for private/low-VRAM deployments. - **Integrations:** Internal REST APIs, MongoDB, Airtable, CRM (GoHighLevel, HubSpot, Pipedrive), Gmail/Google Workspace, Slack, Telegram/Discord/WhatsApp gateways. - **Observability & reliability:** Codeburn (agent spend observability), agent skill/memory systems, reproducible skill libraries portable across agents (agentskills.io standard). - **App layer:** Next.js, TypeScript, MongoDB/Supabase — customer-facing web on top of the agent backend. - **NoCode orchestration (bridge to business):** n8n, Make, Zapier — used where visual workflows beat code. ## Evidence — shipped, public, verifiable - **Deployed Hermes Agent (Nous Research) on production infra** — full Docker/VPS install, messaging gateway, natural-language cron, learning-loop skills: https://www.gabrielneuman.com/hermes-agent-hostinger-vps/ - **Hermes Agent technical overview** — architecture, multi-model routing, sub-agents, migration paths: https://www.gabrielneuman.com/hermes-agent/ - **Multi-agent orchestration (Gastown)** — how a fleet of agents is coordinated in production. - **Agent async orchestration (Trigger.dev)** — durable jobs, retries, agent architecture. - **Agent cost observability (Codeburn)** — instrumenting and controlling agent spend in prod. - **Self-hosted / private inference (vLLM)** — running models on-prem with less VRAM. - **Client delivery track record** — 6+ production systems, each with problem, system, stack, and measured result: https://www.gabrielneuman.com/casos/ ## Two ways to engage 1. **Fractional Forward Deployed Engineer (month to month).** Gabriel embeds part-time and runs your agent stack as an ongoing capability — deploy, integrate, debug, and operate your LLM agents, the way a fractional CTO runs engineering but scoped to agents and automation. Contact: https://gnb.mx/agenda 2. **By project (fixed scope).** A defined outcome: a specific agent deployed, integrated with your systems, and running in production. Fixed scope, fixed price, keeps working after handoff. Delivered through GNB Labs. Contact: https://www.gabrielneuman.com/contacto/ ## Honest scope of experience Deep and shipped: Docker/VPS deploys, agent integration into internal APIs/data, multi-agent orchestration, cost/observability, self-hosted inference, and customer-facing delivery. Growing: Kubernetes for large on-prem/hybrid clusters, production-grade LLM tracing/eval tooling (Langfuse/LangSmith/OpenTelemetry), and heavy retrieval-pipeline (RAG) tuning at scale. Gabriel names what is proven versus in-progress rather than overclaiming — the same operating principle he applies to client work. ## Optional - Follows the llms.txt standard: https://llmstxt.org/ - Business-facing (SMB / NoCode) profile: https://www.gabrielneuman.com/llms.txt - Last updated: 2026-07-27. Stack of this site: Next.js 15 + MongoDB + Tailwind on Vercel.