Have questions about how the Nutanix Platform works? Looking to get started - start here!
Recently active
Failures are part of everything and Nutanix Clusters is not immune to it. But how we plan for failures determines the versatility of the product or a person for that matter!!Nutanix categorizes the type of failures into availability domains essentially based on type of failure. Nutanix provides the ability to tolerate rack failure for extended data availability, in addition to drive, node, block and network link failure. Node FailureA Nutanix Node comprises Physical host and a controller VM. Both these components can fail without any impact to the Nutanix cluster.CVM failureWhen a CVM fails, an alert is generated in Prism and another CVM redirects the storage path on the related host to another CVM. Read and writes will occur over the 10GbE network until the CVM comes back online.It is business as usual for the end customer with maybe a slight performance decrease.Controller VM FailurePhysical Host failureIf a node fails, all HA-protected VMs can be automatically restarted on other nod
Hi folks, In VMware, we need to enable EVC to be able to do vMotion between diff processors, what about AHV? what should be done or is it working smoothly?
Below are new knowledge base articles published on the week of March 21-27, 2021.KB 10516 - [ Karbon ] PE cluster is showing alerts for VGs used by the Kubernetes cluster(s) KB 10651 - NCC Health Check: metering_rest_connection_check KB 10768 - A number of VMs may be missing from the list when monitoring cluster using SNMP protocol KB 10813 - "UnicodeDecodeError" and "UnicodeEncodeError" for VM operation KB 10936 - Duplicate scheduled reports triggered after Daylight Savings Time (DST) change KB 10946 - Identifying the source IP generating TCP Reset packets in a network path KB 10953 - Using Nutanix Objects Self-Signed Certificate with Veritas Enterprise Vault KB 10954 - SMCIPMITool commands output "The node product key needs to be activated for this device" on BMC 7.10 KB 10967 - Cloning a Secure boot enabled VM on AHV with the "Custom Script" option enabled fails with "q35 machine type does not support ide bus type" error KB 10976 - Cluster instability after upgrading both primary an
NVIDIA GPUs primarily have two modes of operation: Compute and Graphics.Compute Mode: the GPU operates within a configuration that is optimized for high-performance computing applications.Graphics Mode: the GPU is optimized for graphics processing and can subsequently be assigned into vGPU profiles for virtual machines (vGPU profiles cannot be used while in compute mode).Various NVIDIA GPUs are provided with default configurations for either of these modes and, sometimes, it is necessary to change the mode to better suit the corresponding workload of the host.In previous models of GPUs, it has been necessary to temporarily boot an AHV host into a NVIDIA-provided Linux ISO and invoke a “gpumodeswich” command with options to apply this change. With newer models of GPU, a command can be found natively within the AHV host filesystem after the corresponding GRID driver has been installed.You can find more information regarding this command via the “Nvidia: Unable to Assign vGPUs to guests w
We have numerous clusters and each have their own hardware platform (and in a couple of cases, we’ve even mixed hardware models within the same cluster). Our current Nutanix footprint;VENDOR MODELS NODE COUNTSuper Micro NX-8155-G7 12Super Micro NX-1175S-G6 7Super Micro NX-3060-G7 6Dell XC630-10 14Super Micro NX-8035-G7 12Super Micro NX-8035-G6 36Is there a preferred or recommended solution for hardware monitoring and alerting? For instance; Dell OpenManage (DOM) or Supermicro Server Manager (SSM). I don’t readily know if either would support the others solution (like DOM supporting IPMIs or SSM supporting iDracs) Or, would Prism Central (and Prism Element) be sufficient for hardware monitoring and alerting?We currently have and are using DOM but not for the SuperMicro hardware so I’m wondering if I’m potentially missing out on a better solution or not.
Nutanix Era is a suite of software which automates and simplifies database management, bringing one-click simplicity and invisible operations to database provisioning and lifecycle management (LCM). Starting with Copy Data Management (CDM) as its first offering, Nutanix Era enables database admins to provision, clone, refresh and restore their databases to any point in time. Through a rich, but simple to use, UI and CLI, they can restore to the latest application-consistent transactionEra enables you to easily provision database environments (either production or otherwise) on your Nutanix clusters. Also, you can only provision the database server VM that hosts a database, so that you can later create or clone databases on that database server.Some of the components include Database engines: Custom software images that are tailor-made to enterprise needs. Database profiles: Customizable database profiles for software, compute, networking, and database parameters. Database recovery
Below are new knowledge base articles published on the week of March 14-20, 2021.KB 9365 - Alert - A802002 - AncDnsUnresolvable KB 10693 - Alert - A130334 - NGT CD-ROM not Unmounted on the VM KB 10694 - Alert - A130192 - Conflicting NGT policies KB 10746 - AHV Guest VM Boot Order Being Modified to Default by Prism After Any Config Save KB 10847 - LCM Pre-check: "test_ncc_checks" KB 10857 - UVMs with VLAN tag may disconnect from network when uplink bond is tagged with VLAN on AHV host KB 10874 - VM migration tasks stuck and libvirt in inconsistent state on clusters running AHV 20170830.x KB 10897 - Xi Frame - Enterprise profile disks growing on every reboot (not extending their partition) KB 10912 - Nutanix Files: Managing Files-At-Root on NFS distributed shares KB 10944 - Alert - A130103 - NGT Mount failed KB 10957 - AHV networking interfaces are renamed after AHV upgrade to 20201105.1082 or later if RDMA NICs are presentNote: You may need to log in to the Support Portal to view some o
A single-node cluster is configured like a regular (three-node or more) cluster in many ways, but here are some of the conditions. Nutanix offers the option of a single-node cluster for ROBO implementations and other situations that require a lower cost option and accept lowered resiliency protections. Single-node clusters are supported only on a selected set of hardware models. Refer the following article for details single-node-supported-hardwares Do not exceed a maximum of 1000 IOPS Do not exceed a maximum of 5 guest VMs . To protect the guest VMs from a scenario of node failure, nutanix recommends to configure backups. These are unlike single-node replication targets which are for replication and backup purposes. LCM is supported for software updates, but not firmware updates. There is no built-in resiliency for Prism Central on a single-node cluster. Do not create a Prism Central instance (VM) in the single node cluster. Async DR is supported for 6 hour RPO only Use
From the Nutanix Bible: “when a Curator full scan runs, it will find eligible extent groups which are available to become encoded. Eligible extent groups must be "write-cold" meaning they haven't been written to for awhile”.I have 2 questions:if I have some data just written in the oplog, how long takes to copy those datas in the extent store, assuming that there is not any read of it?Once the data are in the extent store, how long it takes before the erasure coding is applied to the datas?
Is there a command to update a vm’s custom script/user data?I tried the following and it doesn’t work:<acropolis> vm.update my-vm cloudinit_userdata_path=file:///tmp/user-data.txt
Use the Data Services IP method for external host connectivity to VGs. For backward compatibility, you can upgrade existing environments non disruptively and continue to use MPIO for load balancing and path resiliency. For security, use at least one-way CHAP. Leave ADS enabled. (Enabled is the default setting.) Use multiple disks rather than a single large disk for an application. Consider using a minimum of one disk per Nutanix node to distribute the workload across all nodes in a cluster. Multiple disks per Nutanix node may also improve an application’s performance. For performance-intensive environments, we recommend using between four and eight disks per CVM for a given workload. Use dedicated network interfaces for iSCSI traffic in your hosts. Place hosts that use Nutanix Volumes on the same subnet as the iSCSI data services IP. Use a single subnet (broadcast domain) for iSCSI traffic. Avoid routing between the client initiators and CVM targets. Receive-side sca
Are there ANY plans to expand on the PowerShell cmdlets at all, specifically for getting host info? They’re pretty bare bones. I don’t even know the last time the cmdlets were updated/expanded. It seems to me Nutanix regrets making them and hopes they will quietly die. I have relied heavily on PowerShell to keep tabs on a very large Nutanix environment running Nutanix on both AHV and ESXi. I’m talking 700+ nodes. The ESXi clusters are no problem getting the info from because I use VMware’s PowerCLI to supplement the lack of data I can get with Nutanix Cmdlets, but the AHV clusters are another story since I don’t have PowerCLI to fall back on. Can we at least get a cmdlet that will allow us to use PowerShell to initiate ncli or acli commands, like VMware does with the “Get-EsxCLI” cmdlet. That would be very helpful.
I’m sure you have seen that one before. In most cases you expect it or at least understand what caused it. In some instances you probably ignore it (we all do, no shame). What if this happens when you log into the CVM or the host? Has cluster security been compromised?During the upgrade or rescue of the AOS new keys are created for each node in the cluster. When you open SSH session, these keys are compared to those that were noted on the client previously and since there is a mismatch a warning is triggered.KB-2388 Upgrade/Re-install of AOS changes the ssh key for remote host identification explains how to clean up the keys to get rid of the warnings.
Hi Community,Version and model - NX-1065-G6I have a faulty DIMM in one node. Memory | Uncorrectable ECC (@DIMMC1(CPU2)) | AssertedI would like to remove it from the node and restart it. i.e. Not to replace it with a new memory module.Is this possible? Risks? How should I configure the node with the missing DIMM, or is it something that Nutanix takes care of automatically? Any other advice? Recommendations?I found this document. Is this the correct one to follow?Cheers!
Hi,We are planning to move to Nutanix in our organization. One of the application that is in the scope of this project is Splunk. We are a small environment and our Splunk data intake is around 30GB/day and the setup is mainly used as a SIEM.We have received a few recommendation to isolate Splunk to a separate cluster. Is this necessary, or can we have it on the same cluster if we could guarantee the availability of resources for it? The cluster will host a few application and some infrastructure component like AD and DNS.Are there any major benefits if we isolate Splunk in a different cluster?
A Storage Replication Adapters (SRA) allows VMware Site Recovery Manager (SRM) to integrate with 3rd party storage array technology. Nutanix SRA is one such software module construct that allows SRM to interact with Nutanix clusters. This allows SRM to perform “Array Based Replication” using Nutanix replication, Data Protection and disaster recovery.SRM is an orchestration tool that allows us to perform recovery plans and runbook functionality. You can set up, test and perform pre and post recovery steps in case of a failover, Set the order in which the VMs come up and decide to change IP address if needed. SRM depends on vCenter and is licensed separately. Before you start, ensure the following. SRM and vCenter versions are compatible. Please visit the VMWare website to confirm compatibility. AOS SRA and SRM versions are compatible. Please refer to the Nutanix portal Nutanix SRA for SRM compatibility matrix to confirm compatibility. You will need 2 clusters managed by 2 vCenter
Below are new knowledge base articles published on the week of March 7-13, 2021.KB 9484 - Alert - A802003 - VpcRerouteRoutingPolicyInactive KB 10545 - Alert - A110457 - StaleVMPresent KB 10898 - Objects - Unable to register s3 endpoints that are self signed | Peer certificate cannot be authenticated with given CA Certificates KB 10909 - Move fails with SCSI VirtIO device driver not found if RedHat includes kernel version 2.6.32-220 or prior in /boot.Note: You may need to log in to the Support Portal to view some of these articles.
I wanted to better understand syslog events for a given AOS cluster. It appears that a single node is designated as the ‘syslog leader’ and forwards all events to the destination collector. Thus, the remaining nodes send little to no events to the collector. Is this correct?My clusters run AOS version 5.15.4 LTS for what it’s worth.
Nutanix Insights is a new software-as-a-service (SaaS) offering that aims to redefine the Support experience for our customers, and significantly improve the health of their clusters, by leveraging the telemetry we receive from clusters where a customer has activated Pulse. Nutanix includes a set of features known collectively as Insights that provides a predictive health and support automation platform. Insights dynamically analyzes the extent to which you are following best practices in configuring your clusters for long-term reliability, availability, and performance.Insights works as follows: Pulse collects cluster data and sends it to Nutanix customer support. The Pulse data goes to the Insights engine, a SaaS-like service in the cloud, that does deep analytical processing of the Pulse telemetry and identifies potential issues based on findings or patterns in the data. Insights employs analytics built on historical data and best practices to identify cluster configuration gaps t
Acropolis Dynamic Scheduling proactively monitors the nutanix cluster for any compute and storage I/O contentions or hotspots over a period of time. If ADS detects a problem, ADS creates a migration plan that eliminates hotspots in the cluster and migrates VMs from one host to another. We can monitor VM migration tasks from the Task dashboard of the Prism Element web console Following are the advantages of ADS ADS improves the initial placement of the VMs depending on the VM configuration Nutanix volumes uses ADS for balancing sessions of the externally available iSCSI targets By default, ADS is enabled and Nutanix recommends to keep this feature enabled. ADS monitors the following features: VM CPU Utilization: Total CPU usage of each guest VM. Storage CPU Utilization: Storage Controller (Stargate) CPU usage per VM or iSCSI target. ADS does not monitor memory and networking usage Lazan is the ADS service in an AHV cluster. AOS selects a Lazan manager and Lazan solver among the h
Just a few questions re sequential I/O and the OPLOG versus the Extent Store. If the I/O is deemed sequential in nature, will this always bypass the OPLOG or only when the write operation is larger then 1MB? Does bypassing the oplog mean that the write will be a lot slower in comparison? It still hits the SSD so my assumption is that it’s going to be the same. Why does the process of coalescing the writes before sequentially draining them help with performance? I’m interested in why this step is necessary as opposed to just writing directly to the SSD and then replicating out.
Hi,I can’t access LCM in my CE cluster. It just says “LCM Framework Update is in progress. Please check back when the update process is completed.” It’s been like this for days. Updating Foundation and NCC worked fine the old way.I can’t see any active tasks in Prism or using ecli task.listAny ideas?
I am interested on visualizing east west traffic as well as egress within my NTX environment. With Open vSwitch it seems to be supported through the default vTap interface on the bridges created on my interfaces. My question is, has anyone successfully set up a Nutanix vTap interface to capture packet flow on an external appliance and are there any negative side effects of such a design (overhead, congestion, latency… etc)? I do have an upstream appliance that can handle the analysis similar to Gigamon, CloudLens and ExtraHop. https://docs.openvswitch.org/en/latest/topics/tracing/
Below are new knowledge base articles published on the week of February 28-March 6, 2021.KB 10526 - Alert - A200401 - VM forcibly powered off KB 10856 - Cluster Maintenance Utility version via cli KB 10861 - Acropolis crashing with "KeyError: u'off'" error after upgrade to 5.19.x KB 10863 - Non-admin AD users cannot Clone a VM when added to a Custom RBAC role on PC KB 10878 - VMSA-2021-0002 / ESXi OpenSLP / Disable CIM Server KB 10882 - [Nutanix Move] 1-click-upgrade from 3.7.1 to 3.7.2 may fail when we have migration plans which were created in Move version less than 3.7.0Note: You may need to log in to the Support Portal to view some of these articles.
The flash mode for VM allows to set the storage tier preference to SSD for a virtual machine or a volume group. Without flash mode, data for a mission critical application such as a relational database can run out of the room in the SSD tier because other workloads running on the same cluster. When this happens, the database could potentially migrate to the HDD tier. For extreme latency sensitive workloads, this migration to the HDD tier could negatively affect the read and write performance. By default, you can use up to 25% of the cluster-wide SSD tier as flash mode space for VMs or VGs. If the data size for flash mode enabled VMs or VGs exceeds 25% of the SSD capacity, the system may down migrate the data. Before down migration, the flash mode feature tries to preserve the excess data on the SSD tier for some reasonable amount of time so that you can take corrective actions on the cluster and bring it back to stable. To reduce flash mode usage, we can disable the flash m
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.