Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
Let's say for example we have multiple VLANs in our environment to logically segment the traffic. We need to configure the network interface accordingly as we might need a VM to be VLAN aware. Before understanding the method to configure this, let us understand the difference between trunk and access modes. An access port sends and receives untagged frames (i.e. all frames are in the same VLAN) A trunk port supports tagged frames and thus it allows to switch multiple VLANs. Do we have a method to configure trunk mode in a NIC The following article mentions the steps to safely change a NIC mode of a VM to trunk mode. How to change NIC mode (Access, Trunked)
Hello, Which 10G Switches and SFP+ modules are supported with NX-3360-G7 nodes & C-XCVR-SR-SFP+ modules?
I have two upgrade options showing in my Nutanix clusters. One is for a 5.15 release and the other is for a 5.16 release. I'm currently on a 5.11.2.3 release and generally keep my clusters updated to the latest release possible. Normally I'd go to the 5.16 release and not ask this question, but I heard from someone that 5.16 might have some problems in it. Would I be safe upgrading to the 5.16 release, or would it introduce instability into my environment?
Pulse is an essential tool for maintaining uptime on a Nutanix cluster. While alert emails can directly open a case for an issue which has already happened, the data gathered and sent by Pulse enables identification of potential known issues that haven’t impacted your cluster yet. Enabling this feature within a Nutanix cluster is fairly simple, but depending on your network setup and security there may be some additional steps to make sure it’s working. When you first set up your cluster, right around the time you accept the EULA and change the password for ‘admin’ you are given the option to disable Pulse. So long as you don’t select to disable it, Pulse will attempt to work with the default settings. For some environments it’s just that simple. The cluster will start sending data periodically to Nutanix and our Support Portal will highlight any concerns identified based on that configuration. The configuration is once-per-cluster. If you want to check or update your Pulse configu
Where we can see and create webhooks to integrate with third party monitoring tools e.g. BigPanda.
The NCC health check duplicate_hypervisor_ip_check detects IP addresses that conflict with any of the PE/clusters hypervisor host IPs on the same network. It does so by enumerating all available IPs for the external hypervisor interfaces in the cluster, checks responses to these IPs on the local network, and calls out any duplicate IP's which may be detected. This check results in a PASS status if none of the Hypervisor IPs have been duplicated in the network. If this check returns a FAIL status, that means that either the Hypervisor External IP or Hypervisor backplane network has a conflicting IP address assigned with another device on the same network. What can be the impact? Hypervisor host connectivity can become unstable or unavailable, leading to performance impact, redundancy concerns, and potential downtime. Below is an overview of the steps involved in case the check reports a failure: Run an arping command from any of the CVM and ensure a reply is received from onl
In a traditional Nutanix cluster at least 3 nodes are expected to form a cluster. There is an option however to form a single-node or a two-nodes cluster for ROBO (Remote Office/Branch Office) implementations or as a backup site. The working of the two-node cluster is different from our usual 3 or more nodes cluster. Some examples are: We cannot expand a two-node cluster. Node removal is not supported. There is no cluster stop for 2 node cluster So what is the graceful way to shut down and start a two-node cluster? How to shut down a 2-node cluster: Ensure that the cluster has data resiliency OK and can tolerate one node down from Prism. Stop user VMs - graceful shutdown. There is no cluster stop for 2 node clusters. Log in to a CVM using the nutanix account, and perform a graceful shut down of the CVM. Wait for 5-10 mins, and then shut down the second CVM. Shut down the hosts. NOTE: The above two commands used to shutdown the CVMs are different. Take a look at
While attempting to update a VM category via the prism central API, we are getting a 409 response saying: { 'api_version': '3.1', 'code': 409, 'message_list': [{ 'message': 'Edit conflict: please retry change. Entity CAS version mismatch.', 'reason': 'CONCURRENT_REQUESTS_NOT_ALLOWED' }], 'state': 'ERROR'} It seems to be occurring on only one VM as all others seem to update normally. I can clear it by updating the VM through the gui, however once I try and do it via API I get that error. Any thoughts on what could be going on with this VM?
Are you looking for a way to support more than 12K VMs as part of your Prism Central installation? Then you should try to scale out Prism Central. Nutanix introduced Prism Central scale-out architecture in AOS 5.6. Prism Central scale-out architecture allows customers to scale out its Prism Central deployments incrementally, depending on the needs. Single Prism Central instance can support up to 12k VMs, with Prism Central scale-out, the number of supported VMs increases up to 25k PoweredOn VMs. Prism Central scales-out has 3 VMs and has been architected to tolerate one node failure (n+1 fault tolerance). The following requirements must be met before you can expand Prism Central or deploy a Prism Central VM: The specified gateway must be reachable. No duplicate IP addresses can be used. The container used for deployment is mounted on the hypervisor hosts. When installing on an ESXi cluster: vCenter and the ESXi cluster must be configured properly. See the vSphere Administra
Flow is a software-defined networking product tightly integrated into Nutanix AHV and Prism. Flow provides rich visualization, automation, and security for VMs running on AHV. There are instances in which flow is not configured in your environment but still, you will see the following alert "Flow Control Plane Failed" There could be multiple reasons for the alert to be recurring but if you're sure that you have not enabled flow in your infrastructure, you can use the following KB article to get a better idea about the alert, confirm if the flow is enabled or not and the possible workarounds "Flow Control Plane Failed" alert appears on PE even if Flow is not enabled Want to know more about Flow and it's best practices? The following document will help you understand the architecture and common guidelines regarding Nutanix Flow Nutanix Flow Guide
Hi everyone, I’m looking for your advice please for ESXi-based Nutanix clusters with hybrid storage. Question is regarding the placement of ESXi persistent scratch. With traditional architecture, the advice normally was to configure /scratch to be on a shared datastore, not a local one. The argument was that in case of a host failure, the logs will still be available from other hosts. With Nutanix, I can see the options as 1) leave the default, on vfat partition on SATADOM/M2/BOSS; 2) in a folder on local VMFS datastore where CVM is (same storage device though); 3) on a shared datastore, i.e. on DSF. What is the Community (best) practice with ESXi /scratch in Nutanix environment? Please kindly share your advice and thoughts. Thank you.
Two weeks ago I’ve upgraded our two clusters to AOS 5.15 . After that, both clusters began to receive constant warnings for root partiton space usage high(exceeded 80%) on almost all CVMs. Below is the result of ‘df -h’ on one of the CVMs: Filesystem Size Used Avail Use% Mounted on devtmpfs 18G 0 18G 0% /dev tmpfs 512M 0 512M 0% /dev/shm tmpfs 18G 1.2M 18G 1% /run tmpfs 18G 0 18G 0% /sys/fs/cgroup /dev/md1 9.8G 7.4G 1.9G 80% / /dev/loop0 240M 2.3M 221M 2% /tmp /dev/md2 40G 23G 17G 57% /home tmpfs 3.6G 0 3.6G 0% /run/user/1000 /dev/sdg1 5.5T 1.4T 4.0T 26% /home/nutanix/data/stargate-storage/disks/ 17Q0A01WFB9D /dev/sdh1
I am trying to locate the memory slots “P2-DIMMD1” and “P2-DIMMD2” in G4 but did not see any slot marked as P2-DIMMD1 or P2-DIMMD2. would you please let me know how to find these slots and replace memory. Thanks, Tairshah
Our vision at Nutanix has been and will always be: "one platform, any app, any location". This has been our goal from close to the beginning. We have openly documented Nutanix architecture in the freely available Nutanix Bible, and we are committed to open-source software, actively using and contributing code within a variety of communities. For example, here is a simplified architecture of a Nutanix environment taken from the Nutanix bible: In the documentation, you will find all the information you need regarding the Software, Hardware, performance, recommended practices, restrictions and more.
Below are new knowledge base articles published on the week of April 12-18, 2020. KB 9109 - Cannot Login to Prism with AD account KB 9168 - Custom Virtual Machine Report Creation in Prism Central KB 9223 - Nutanix Files: FileServer preupgrade check failed with cause(s) Sub task poll timed out KB 9224 - Restore the SSR snapshots using CLI KB 9232 - LCM Failure: Could not find pnics attached to the CVM interface KB 9239 - LCM upgrade failure on HPE - The node is not in production mode KB 9264 - AHV IDE bus performance implications Note: You may need to log in to the Support Portal to view some of these articles.
I am deploying nutanix at one of our customers site but unfortunately now he has 2 Switches with 1 GIG ports so is it doable to remove the 10 GIG from the bond and rely only on the 2*1G ports ? Model used (NX-1365-G6) thanks in advance
The NCC health check check_vcenter_connection verifies if the vCenter Server is registered with Prism and if a connection can be established. Nutanix cluster communicates with vCenter Server to obtain virtual machine information necessary for certain Nutanix cluster operations like Data Protection, One-Click upgrades, etc. If the vCenter Server is not registered or is not accessible, those operations may fail. The check returns a PASS if vCenter Server is registered with Prism Element and connection can be established. The check returns an INFO if vCenter is not registered with Prism The check returns a FAIL if vCenter is registered with Prism and the connection to vCenter cannot be established. This scheduled to run every 5 minutes, by default and will generate an alert after 3 consecutive failures across scheduled intervals. To take a look at the NCC check and the solution section https://portal.nutanix.com/page/documents/kbs/details/?targetId=kA032000000TVQACA4 For instructio
To perform core VM management operations directly from Prism without switching to vCenter Server, you need to register your cluster with the vCenter Server. Nutanix cluster communicates with vCenter Server to obtain virtual machine information necessary for certain Nutanix cluster operations like Data Protection, One-Click upgrades, etc. If the vCenter Server is not registered or is not accessible, those operations may fail. The NCC health check check_vcenter_connection is also in place to verify if the vCenter Server is registered with Prism and if a connection can be established. Follow the steps to register your cluster with vCenter: Log into the Prism web console. Click the gear icon in the main menu and then select vCenter Registration in the Settings page. Click the Register link. Enter the administrator user name and password of the vCenter Server in the Admin Username and Admin Password fields. Click Register. Following are some of the important points about regi
Did you know that you can directly access the files in the Nutanix container from your local desktop? It is possible using WinSCP. This process helps in uploading or downloading files to and from Nutanix container. For example, you want to download a phoenix iso you generated on a CVM. Here are the steps to access a container: Open WinSCP. Connect to the CVM IP using SFTP protocol and port 2222. Login using the admin/prism element credentials. Enable the option to show hidden files by going to Options > Preferences > Panels and then selecting the “Show hidden files” option under the common settings. From here you can either upload or download files to the container. Note: Do not delete any data from the container via WinSCP or similar tool. Appropriate Prism or CVM command-line workflows should be leveraged to perform the cleanup if needed. To take a look at the steps in detail, take a look at https://portal.nutanix.com/page/documents/kbs/details/?targetId=kA0
We all know that Virtual IP is used to access the Prism web console. It is also referred to as the external IP address of the cluster. Cluster virtual IP is mapped to the CVM which is the Prism service leader. Every time a new leader is elected the virtual IP is transferred to the new leader CVM, ensuring Prism Element availability. The NCC health check virtual_ip_check verifies if the cluster virtual IP is configured and reachable. It is scheduled to run every hour, by default and will generate an alert after 1 failure. To manually verify virtual IP settings in the Prism web console. Click the cluster name in the main menu of the Prism web console dashboard. In the Cluster Details pop-up window, check that a Cluster Virtual IP is configured and is correct. Check if the virtual IP configured is in the same subnet as the CVM IP. The settings can also be accessed from CLI. Have a read about this check and the scenarios where it can throw alerts https://portal.nutanix.c
Did you know that in Acropolis (AHV) you can enable high availability for the cluster to ensure that VMs can be migrated and restarted to another node in case of a failure? Best effort VM availability is enabled by default in Acropolis. Virtual Machine High Availability (VM HA) VM HA is a feature designed to ensure that critical VMs are restarted on another Acropolis Hypervisor (AHV) host within the cluster if a host fails. There are two VM high availability modes: Default - This does not require any configuration and is included by default when an Acropolis Hypervisor-based Nutanix cluster is installed. When an AHV host becomes unavailable, VMs that were running on the failed AHV host are restarted on the remaining hosts, based on the available resources. Not all of the failed VMs will restart if the remaining hosts do not have sufficient resources. Guarantee - This non-default configuration reserves space to guarantee that all failed VMs will restart on other hosts of the clu
Let’s say you need to administer your user VMs from the command line interface. The Controller VM resources are shown under the VM page In the Nutanix Prism, but you will not be able to change the resources configuration unless you connected to the Acropolis hypervisor (host) and modified the configurations using virsh. “virsh: is a command line interface tool for managing guests and the hypervisor.” Centos.org. First you can review the settings of the CVM under the VM Page on Prism. Connect to the Acropolis hypervisor (host) using the root account with password “nutanix/4u” Lists all the VMs on a host > virsh list –all Displays information about a VM > virsh dominfo VM_Name Displays information about the vCPU > virsh vcpuinfo VM_Name Sets the number of virtual processors > virsh setvcpus VM_name count Note: The count value cannot exceed the number of processors specified for the guest. You can increase the number of processors by editing the virs
Let's say that you ran the health checks on your cluster and received data_replication_check failure, what does it mean and how do you fix it? The NCC health check data_replication_check helps to ensure that the customers are not impacted by an extremely rare condition which can result in the inability to restore from snapshots. This issue is covered in Field Advisory 28. Scenario 1 - Cluster is running an old version of AOS.Scenario 2 - Cluster was recently upgraded, but has snapshots created before the AOS upgrade. For a full explanation and the solution for both scenarios check out the health check documentation at: KB-2089. To make sure that the issue is resolved you can run the dedicated health check for this component by connecting to a CVM and running: "ncc health_checks data_protection_checks protection_domain_checks data_replication_check" Note: This health check has been retired from NCC 3.9.3.
Scalability is the backbone of any industry solution, hence with ever-increasing workloads, it is natural for us to add new nodes in our existing infrastructure to make it scalable and highly efficient. Nutanix provides a feature which allows you to add new nodes to your existing cluster and increase the cluster overall capacity. A question lingering in our mind right now Do we have some guidelines regarding cluster expansion? Cluster expansion depends on the AOS version, hypervisor type (AHV, ESXi, Hyper-V, or Citrix Hypervisor), data-at-rest encryption status, and certain hardware configuration factors. The following documents provides some basic guidelines specific for cluster expansion Cluster Expansion .NEXT:Want to expand your cluster
It may be helpful to determine what commands were run previously to troubleshoot a current issue. Let’s first understand what aCLIand nCLI command utilities are. aCLI: utility to create, modify and manage VMs in AHV. nCLI: utility to manage cluster operations. The history files are hidden and are persistent across reboots. Below article explains how to retrieve commands history: https://portal.nutanix.com/page/documents/kbs/details/?targetId=kA032000000TSsjCAG Also, take a look at how to check cluster upgrade history: https://next.nutanix.com/discussion-forum-14/checking-cluster-upgrade-history-37417
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.