
A self-hosted or cloud platform for Kubernetes and multi-cloud operations: AI assistants that triage alerts, investigate incidents to a root cause, and cut cloud costs, plus a no-code automation builder to run the fixes. AWS, Azure and GCP.
A startup from Pune, India that is founded by Rakesh Rajendran, Shiv Pratap Singh.
This page is designed to help you find out whether NudgeBee is good and if it is the right choice for you.
NudgeBee is a CloudOps platform for teams running Kubernetes, cloud infrastructure on AWS, Azure and GCP, or both. Connect a cloud account, a cluster, or any combination. It builds a live topology map of the estate and turns a stream of disconnected alerts into a short list of ranked, correlated issues.
Four assistants share one backend:
Diagnostics run read-only, and every create, update and delete is gated behind explicit human approval. Once approved, it remediates: a cloud API change, a live Kubernetes patch, an Auto Pilot rightsizing run with guardrails, or a ready-to-merge pull request that traces the failure to the exact source line, commit and author.
Investigations reach people where they already work. Ask Nubi in a Slack, Microsoft Teams, Google Chat or Discord thread, or by email, approve or reject the fix from buttons in the message, and have the root cause written back onto the PagerDuty or Jira incident.
No model lock-in: AWS Bedrock, OpenAI, Azure OpenAI, Vertex AI, Anthropic, Ollama and vLLM, including self-hosted endpoints.
Run it in your own cluster with a Helm chart, or use the hosted cloud. Works with the stack you already run: Prometheus, Grafana, Datadog, Jaeger, Jira, ServiceNow, GitHub and GitLab.
Listed in
Deployment
Self-hosted in your own Kubernetes cluster. Helm umbrella chart, with AWS, Azure and GCP Terraform modules.
Licensing
Source-available. Self-hosting from source is free and unlimited for your own internal use.
Cloud Compatibility
AWS, Azure and GCP in one collector, one data model, one recommendation taxonomy.
Four Assistants, One Backend
SRE, FinOps, K8s Ops and CloudOps as four personas on one shared platform, not four separate tools.
Root Cause Analysis
Mandatory five-why chain with cited tool evidence. Symptom-only answers such as 404, CrashLoopBackOff or "resource missing" are rejected by a final-answer critiquer.
Causal Correlation
Root cause separated from symptom by graph traversal, not time-window grouping.
Human Approval
Investigate, do not execute. Every create, update and delete is classified and gated behind explicit approval, enforced in prompts and in the tool-access layer.
Bring your own model
11 LLM provider routes: Bedrock, OpenAI, Azure OpenAI, Google AI, Vertex AI, SageMaker, HuggingFace, Anthropic, Ollama and vLLM.
Kubernetes Operations
Live topology of workloads, pods, nodes and namespaces, with in-browser pod exec, logs, metrics, traces and profiler.
Event De-noising
Owner-level dedup. Crash loops, image-pull backoff, OOM kills, job failures and CPU throttling collapse to one issue per workload.
Observability Integrations
Queries 19+ existing backends in place via native dialects including PromQL, LogQL, KQL, NRQL, Grail-DQL, SignalFlow, OPAL and ES-DSL. No rip and replace.
Anomaly Detection
Three swappable ML engines (IsolationForest, DBSCAN, Z-score) across CPU, memory, latency, error rate and replicas.
Knowledge Graph
Multi-cloud topology graph with behavioural edges from five runtime sources including eBPF and traces. PostgreSQL only, no graph database.
FinOps Rules
Provider recommendation rules across AWS, Azure and GCP in four categories and five severities, each one dollar-quantified.
Kubernetes Rightsizing
Vertical (CPU p99, memory peak plus 15 percent, OOMKill-aware), horizontal replica forecasting, node-fleet optimization by integer linear program, PVC and spot migration.
Savings Prioritization
FinOps Score 0 to 100 with Act Now, Critical, High, Medium and Low bands, recomputed every six hours.
Cloud-Native Savings
Folds AWS Compute Optimizer, Cost Optimization Hub, Cost Explorer, Trusted Advisor, Azure Advisor and GCP Recommender into one dollar-normalized model.
Automation Engine
No-code runbooks executed on Temporal. Durable, crash-safe, resume-exactly, versioned with a live pointer.
Auto Pilot
Scheduled autonomous rightsizing with dry-run, resource filters, guardrails and change-gated notifications.
ChatOps
Conversational in Slack, Microsoft Teams, Google Chat, Discord and email, with interactive approve and reject actions.
ITSM
Jira, ServiceNow, PagerDuty, Zenduty, GitHub Issues and GitLab Issues behind one normalized API, with the root cause written back onto the PagerDuty or Zenduty incident.
Security Scanning
kube-bench (CIS), Trivy CIS and image CVE, Popeye, certificate expiry, version skew, plus ingested AWS GuardDuty, Inspector and Security Hub, and Azure Defender and Sentinel.
Access Control
Eight-tier RBAC from super-admin down to namespace read-only, with Google, Okta, Azure AD, OneLogin, LDAP and magic-link sign-in.
Data Residency
Metrics, logs and traces are queried in place in your own observability backends. They are not shipped to a vendor SaaS.
SRE, Platform Engineering, DevOps, Cloud and FinOps teams running production Kubernetes on AWS, Azure or GCP.
It fits teams that already have observability, ticketing and chat tooling they intend to keep, and want investigation, cost optimization and remediation on top of that stack rather than a replacement for it. Regulated and air-gapped environments are covered by the self-hosted deployment.
Four assistants share one backend: SRE, FinOps, K8s Ops and CloudOps. Most tools in this space do incident investigation only, so cost work and Kubernetes operations end up in separate products with separate data.
It runs in your own cluster from readable source, and self-hosting from source is free and unlimited for your own internal use. There is no model lock-in either, with 11 LLM provider routes including Ollama and vLLM for fully self-hosted inference.
Diagnostics run read-only. Every create, update and delete is classified and gated behind explicit human approval, enforced both in the prompts and in the tool-access layer, so the platform investigates on its own but never changes infrastructure on its own.
Three reasons.
You can read the source and run it yourself, without a per-node or per-seat bill and without your telemetry leaving for a vendor SaaS. Metrics, logs and traces are queried in place in the backends you already run.
One platform covers investigation and cost. An incident traced to an oversized deployment and a rightsizing recommendation for that same workload live in the same topology graph, so the fix follows the finding.
Root cause means root cause. A mandatory five-why chain with cited tool evidence is enforced, and symptom-level answers such as 404, CrashLoopBackOff or "resource missing" are rejected by a final-answer critiquer.
NudgeBee was built for the toil of running production: alert floods that hide the real incident, root causes that take hours of log and dashboard archaeology, and cloud waste nobody has time to chase.
The source was opened in June 2026 under a readable-source licence, so teams can inspect exactly what an autonomous agent is allowed to do inside their infrastructure before they trust it with production. The company is based in Pune, India.
Go, Python and TypeScript across 14 services, deployed by a Helm umbrella chart.
PostgreSQL holds relational state and the topology knowledge graph, with no separate graph database. Temporal runs the durable runbook workflows, RabbitMQ carries the cross-service event bus, Redis handles caching and distributed locks, and Qdrant stores RAG vectors. The dashboard is Next.js. Behavioural telemetry uses eBPF. Terraform modules ship for AWS, Azure and GCP.
We have collected here some useful links to help you find out if NudgeBee is good.
Check the traffic stats of NudgeBee on SimilarWeb. The key metrics to look for are: monthly visits, average visit duration, pages per visit, and traffic by country. Moreoever, check the traffic sources. For example "Direct" traffic is a good sign.
Check the "Domain Rating" of NudgeBee on Ahrefs. The domain rating is a measure of the strength of a website's backlink profile on a scale from 0 to 100. It shows the strength of NudgeBee's backlink profile compared to the other websites. In most cases a domain rating of 60+ is considered good and 70+ is considered very good.
Check the "Domain Authority" of NudgeBee on MOZ. A website's domain authority (DA) is a search engine ranking score that predicts how well a website will rank on search engine result pages (SERPs). It is based on a 100-point logarithmic scale, with higher scores corresponding to a greater likelihood of ranking. This is another useful metric to check if a website is good.
The latest comments about NudgeBee on Reddit. This can help you find out how popualr the product is and what people think about it.
Do you know an article comparing NudgeBee to other products?
Suggest a link to a post with product alternatives.
Is NudgeBee good? This is an informative page that will help you find out. Moreover, you can review and discuss NudgeBee here. The primary details have been verified within the last quarter. So they could be considered up to date. If you think we are missing something, please use the means on this page to comment or suggest changes. All reviews and comments are highly encouranged and appreciated as they help everyone in the community to make an informed choice. Please always be kind and objective when evaluating a product and sharing your opinion.