Have questions about how the Nutanix Platform works? Looking to get started - start here!
Recently active
I have AHV cluster at both DC & DR, which has number of VMs being protected with Protection Domain feature. These VMs are being replicated from DC to DR in scheduled manner. Now as a part of DR drill, I need to run these production VMs from DR for period of time and then again make it functional from DC once drill is over. I am concerned over how the delta changes made during drill with VMs at DR will be again replicated back to DC as Protection Domain replication is uni directional (DC to DR) ? Any recommendations ?
Below are new knowledge base articles published on the week of December 15-21, 2019. KB 8204 - Alert - A1061 - vDisk Block Map Usage High Critical KB 8284 - Alert - A130151 - Two node cluster state change to KB 8640 - Prism one click upgrade : Preupgrade/Upgrade options not available after manually uploading metadata json and upgrade bundle KB 8743 - Alerts relating to IPMI sensors report that the component cannot be monitored or be permanently damaged. KB 8747 - Anonymous IPMI user Note: You may need to log in to the Support Portal to view some of these articles.
In scripts we use for some ESXi hypervisor configurations we utilize “allssh”. The problem I’m running in to, though, is I must run each allssh command individually rather than using a script to run multiple allssh commands. The first line will kick off and before it has a chance to complete the second will kick off. For example, here’s where we configure DNS on the hypervisor through a CVM: allssh ssh root@192.168.5.1 esxcli network ip dns server add --server=10.1.1.1 allssh ssh root@192.168.5.1 esxcli network ip dns server remove --server=8.8.8.8 allssh ssh root@192.168.5.1 esxcli network ip dns server add --server=10.1.1.2 The first will kick off then the second goes prior to the first completing. The script then gets confused and stops while still connected to one of the ESXi hosts. Does anyone know of a way to initiate subsequent allssh commands only after the prior one running has completed? Although not very clean, should I put a sleep command in between each?
Data is everything in this modern IT world and we want to have as much storage capacity as possible in our infrastructure. Let’s say you have a new Nutanix cluster and want to know the recommended maximum storage utilisation capacity of the cluster. NOTE :- We should not try to utilise the cluster to it’s peak storage as we need to have sufficient space available in case a node or a disk fails to rebuild the data. Whenever a disk fails in a Nutanix cluster, the extent groups of that disk needs to copied to another disk to ensure fault tolerance, same goes for a node failure.So how can we calculate the maximum storage utilisation of our cluster and if different scenarios, if we have RF-3 or if we have RF-2? Please go through the KB-1557 to understand the formula to calculate the maximum recommended usage for a cluster.
Hi everyone,I’m using the API to pull VM performance stats, however I’m having trouble interpreting what I’m seeing.For instance, I’m pulling “hypervisor.cpu_ready_time_ppm” for one of my VMs and getting the following output:{ "statsSpecificResponses": [ { "successful": true, "message": null, "startTimeInUsecs": 1576458000000000, "intervalInSecs": 30, "metric": "hypervisor.cpu_ready_time_ppm", "values": [ 108, 97, 144, 107, 89, 92, 78, 74, 47, 49, 90,.... Output truncated for brevityI get that the metric is a percentage but obviously you can’t have 144% of time, so how should I interpret these values?What I really need is a guide and/or reference that explains all of these metrics and how to interpret them.I found the following link but it doesn’t really tell me what I want to know:https://portal.nutanix.com/#/page/docs/details?targetId=Prism-Central-Guide-Prism-v51:mul-alerts-user-created-metrics-r.html
In some cases, you might have to permanently remove a physical node / host from a Nutanix cluster. There are two scenarios in node removal. Permanently Removing an online node Removing an offline / not-responsive node in a 4-node cluster, at least 30% free space must be available to avoid filling any disk beyond 95%. You cannot remove nodes from a 3-node cluster because a minimum of three Zeus nodes are required. Some Points to consider before initiating node removal: Sufficient Disk space available on other nodes in the cluster User Virtual Machine relocation (if required) Any software upgrade should not be running Checklist on verifying cluster health status Data resiliency is “OK” (green) in Prism Run a complete “ncc report” either from prism or CVM cli: ncc health_checks run_all Depending on the size of data, node removal can be lengthy process, which involves relocating data from the node to other healthy nodes in the cluster. Node removal also remo
Below are new knowledge base articles published on the week of December 8-14, 2019. KB 8306 - Pre-Upgrade Check: test_if_any_upgrade_is_runningKB 8461 - LCM Pre-check - test_oneclick_hypervisor_intentKB 8545 - NSX-T Support on Nutanix InfrastructureKB 8608 - Finding the serial ID of a bad HDD or SSDKB 8660 - [DIAL] Conversion stuck: Migrating UVMsKB 8670 - Prism Central: After upgrading to Prism Central 5.10.6 on Hyper-V, "Illegal instruction (core dumped)" message results when running NCC.KB 8673 - AHV | VM power on may fail with NoHostResources error when initiated from Prism UI or acliKB 8674 - 3rd party storage might cause issues on a Nutanix systemKB 8675 - Cannot plug out the Phoenix (or other) ISO from the IPMIKB 8690 - Alert - A160061 - FileServerShareAlmostFullKB 8694 - Alert - A400111 | EpsilonVersionMismatchKB 8701 - Dell XC-Hyper-V LCM Inventory won't recognize the installed PTAgent & iSM versionsKB 8702 - LCM Operation Failed. Reason: Failed to validate update request.
If your hypervisor is ESXi you know how vCenter manages all the Vms in the vmware data center. You can either use vCenter for Vm management or Prism to control to do most of the same. This stated long ago after AOS 5.0 was released. Most core VM management functions like creating, cloning, updating and deleting VMs, and attaching/deleting disks and NICs as well as power operations alongside console access and guest tool management But for most of the above you will need to register the prism with vCenter. There are rules and guidelines and requirements and limitations that you need to be aware of. You can find out about these and how to register/unregister prism element with vCenter by reviewing: https://portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v56:wc-management-multi-hypervisor-prism-c.html One thing to notice is that the network traffic between the Prims and vCenter will be much higher than usual when prism is not register with vcenter. Ask any qu
Nutanix Controller VM (CVM) is the brain of a Nutanix Cluster running the Nutanix AOS software to form a highly-resilient cluster comprised of 3 or more nodes. Nutanix CVM offers rich data services, virtual machine management and hosting services, replication services - while ensuring all data and metadata is checksummed to ensure data consistency and integrity. Nutanix CVM communicates over a IP network to all other CVMs and Hosts in the cluster. Thus, changing the CVM IP Address should be planned carefully with all the following and any environment specific points: CVM and Hypervisor Host are required to be in the same subnet (192.168.5.x) Hypervisor host can be multi-homed, but point - (i) is mandatory IPMI subnet should be reachable to and from the CVM Cluster Virtual IP Address iSCSI Data Services IP (used by Volumes, Files, Objects, Karbon, LEAP) Network Segmentation check (Backplane traffic) Remote Sites and any on-going replication Guest VM downtime (this is required for Re-IP
Shutting down and Restarting a Nutanix Cluster requires some considerations and ensuring proper steps are followed - in order to bring up your VMs & data in a healthy and consistent state. Nutanix is a Hypervisor agnostic platform, it supports AHV, Hyper-V, ESXi and XEN. This makes it all the more important to read the following Nutanix KB, which details the steps required to gracefully shutdown and restart a Nutanix cluster with any of the hypervisors. Nutanix KB : How to Shut Down a Cluster and Start it Again?
VDi (Virtual Desktop Infrastructure) is one of the earliest applications that was used in hyper-convergent systems. The closer the storage to the cpu and memory the better the performance. Citrix director plugin, need to be connected to the nutanix cluster. if the connection fails with a message "Unable to connect to the host”, you can easily address it by: 1- Allow ICMP traffic and make sure port 9440 is also allowed between Citrix Director and Nutanix cluster. 2- Make sure there is no proxy server configured in the browser configuration The above is reported in the Nutanix public KB article: https://portal.nutanix.com/#/page/kbs/details?targetId=kA00e000000LKiBCAW
Hey Community, can someone clarify the compression and de duplication features that come with the starter edition of AOS? I can see on the Nutanix website: https://www.nutanix.com/products/software-options that starter has a check box beside inline compression and inline de duplication, but not besides compression and de duplication, Does this mean with a starter license environments can still take advantage of the storage efficiencies of de duplication and compression (on write)? Is there any solution briefs that go into detail regarding the compression and de duplication features of the starter edition. Best Regards,
The Move 3.0 services are now Dockerised and all Move and Move Agent Services now run as Docker Containers. You maybe running Nutanix Move version 3 and are unable to connect it to hosts that reside on certain internal subnet you have. it may be that local docker container (Docker0) has also assigned to the same subnet. This could be addressed buy the following KB article in your Nutanix Portal: https://portal.nutanix.com/#/page/kbs/details?targetId=kA00e000000PVwkCAG Ask questions about it, if you are concerned.
Someone who can support me, was doing the idrac update (Hyper-V hypervisor) and failed, but now the node is out and can not lift the services, I try to restart it from the console and I get the error message:2019-12-11 14:04:41 INFO zookeeper_session.py:131 cvm_shutdown is attempting to connect to Zookeeper 2019-12-11 14:04:41 WARNING lcm_genesis.py:219 Failed to reach a [localhost] where LCM [LcmFramework.is_lcm_operation_in_progress] is up. Retrying... 2019-12-11 14:04:46 WARNING lcm_genesis.py:219 Failed to reach a [localhost] where LCM [LcmFramework.is_lcm_operation_in_progress] is up. Retrying... 2019-12-11 14:04:51 WARNING lcm_genesis.py:219 Failed to reach a [localhost] where LCM [LcmFramework.is_lcm_operation_in_progress] is up. Retrying... 2019-12-11 14:04:56 WARNING lcm_genesis.py:219 Failed to reach a [localhost] where LCM [LcmFramework.is_lcm_operation_in_progress] is up. Retrying... 2019-12-11 14:05:01 WARNING lcm_genesis.py:219 Failed to reach a [localhost] where LCM [Lcm
I want to Migration VM. Nutanix(AHV) → Another Nutanix(AHV) So. I tried to get the qcow2 file. refs https://virtualife.pro/export-an-nutanix-ahv-vm/ First I tried executed command “acli vm.list” → succeed :) The next time executed command “acli vm.get <VM name>” → nothing happend :( Why? -- Version Nutanix 5.10.5 LTS NCC 3.7.1.2 LCM 2.1.4139
If you are running Nutanix hardware, you may be familiar with accessing the IPMI page to load ISO files and looking into hyperviosr remote console among many other functions. You can reach this page from Hardware link in Prism (click on table tab and select the hyperviosr you are concerned about, the lower left of the screen will provide you with an IP link to the ipmi page) Occasionally we need to be concerned about the hardware related messages the node provide us (Event Log) OR hardware component level health of the node. The “Server Health” tab in ipmi page can provide us with these valuable information: ==
Hi guys, Is there any compatibility requirement about AOS versions between 2 DR sites? I am going to upgrade one site to newer AOS version and worrying about their AOS compatibility
Below are new knowledge base articles published on the week of December 1-7, 2019. KB 8546 - Pre-check: test_nsx_configuration_in_esx_deployments KB 8624 - PulseHD shows RED if dmidecode.exe is missing KB 8631 - Accessing Prism Via Citrix NetScaler (ADC) KB 8669 - Nutanix Files - long filename isn't supported yet. KB 8671 - How to determine which M.2 device failed on the node KB 8680 - Metro - Recovery procedure after two-node down scenarios Note: You may need to log in to the Support Portal to view some of these articles.
Hi all, Local Replication is a process in which multiple copies of data are stored within a storage container. These copies exist for fault tolerance. Snapshots are placed locally on the same cluster as the source VM. Thus, If a physical disk fails, the cluster can recover data from another copy. The cluster manages the replicated data, and the copies are not visible to the user. So, what is the difference the Replication Factor option? Because RF is used too for fault tolerance in case of a physical disk failure (or node, ...) Thanks
A traditional Nutanix cluster requires a minimum of three nodes, but Nutanix also offers the option of a two-node cluster for ROBO implementations and other situations that require a lower cost yet high resiliency option. Unlike a one-node cluster (see Single-Node Clusters), a two-node cluster can still provide many of the resiliency features of a three-node cluster. This is possible by adding an external Witness VM in a separate failure domain to the configuration (see Configuring a Witness (two-node cluster)). Nevertheless, there are some restrictions when employing a two-node cluster. The following links will provide you guide lines and information abut configuring the two node clusters:Two-Node Cluster Guidelines Two-Node ClustersAsk any questions to clarify any concerns about the the two node clusters.
Every once in a while due to network infra structure changes or because you have to physical move the cluster to another location, you may have to modify the Cluster IP. This includes CVM, Hypervisor and IPMI ip addresses, netmask and default gateways Unfortunately this operation requires taking some down time as you will need to stop the cluster for the duration of change. Before you start, you need to: 1- Clearing the external virtual ip address of the cluster , and setting new ip address for it 2- Ensuring that the Ntp and Dns servers of the cluster are reachable from new CVM ip address and if they are going to be different, remove the old addresses and add the new ones 3- Check that all hosts are part of metadata store You need to consider 3 different scenarios: 1- Change the IP addresses of the CVMs in the same subnet. 2- Change the IP addresses of the CVMs to a new or different subnet. 3- Change the IP addresses of the CVMs to a new or different subnet if you are moving the cl
Foundation is how we build and configure Nutanix clusters and many customers would prefer to make use of more advanced network technologies like LACP to improve the cluster performance and provide redundancy. LACP increases bandwidth, provides graceful degradation as failure occurs, and increases availability. It provides network redundancy by load-balancing traffic across all available links. If one of the links fails, the system automatically load-balances traffic across all remaining links. Foundation 4.2 introduces LACP support for the standalone Foundation. For more information about the supported hypervisors and requirements, please see the KB article titled: LACP Support in Foundation.
If you are replicating data through DR (Data Replication page in Prism), then you have setup schedules for snapshots so they can be copied to the remote site on scheduled time. When you select a “snapshot” filed in one of the protection domains you created in the Prism UI, one of the fields is “Reclaimable space”. You may observe an spinning wheel continuously and the word “processing” on this filed for some or all snapshots, but you may also notice that snapshot(s) has already been taken and done. So why the spinning wheel for this filed? This field is lazy-calculated by Curator during full scans and populated afterward, so it takes sometime (may be few hours) to show up. Until Curator finishes calculating the value, the field shows Processing in Prism.
Hi, I have updated my Nutanix (AHV) license and checking license status, i see correct license expiry date applied to my clusters. I however still get alert from “Licese Standy Mode” details below Possible Cause The license file has not been applied after cluster summary file generation. Recommendation Apply a new license. When i try to reapply license by downloading csf file, uploading csf file i get error: You’ve uploaded an older/inactive cluster summary file(CSF). To continue, download the latest CSF from Prism Element or Prism Central Help please
Nutanix AOS offers simplicity in managing traditional complex infrastructure tasks. From Virtual machine management, Storage operations, replication - and of course Cluster software and hardware upgrades. As Infrastructure admins, we are well aware of the operational pain points, when it comes to upgrading: Hypervisor Upgrades Storage OS upgrades Firmware Upgrades Management software upgrades the list goes on… With Nutanix One-Click upgrades, customers can upgrade software components and hardware components easily. Software and Firmware needs to be downloaded from Nutanix repositories - which is why it is important to understand what Network Ports are required to be open or can be opened on demand to check for upgrades. Following KB from Nutanix Portal lists the required network ports for different services and upgrade repos endpoints: Recommendation on Firewall Ports Config
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.