Have questions about how the Nutanix Platform works? Looking to get started - start here!
Recently active
Foundation is how we build and configure Nutanix clusters and many customers would prefer to make use of more advanced network technologies like LACP to improve the cluster performance and provide redundancy. LACP increases bandwidth, provides graceful degradation as failure occurs, and increases availability. It provides network redundancy by load-balancing traffic across all available links. If one of the links fails, the system automatically load-balances traffic across all remaining links. Foundation 4.2 introduces LACP support for the standalone Foundation. For more information about the supported hypervisors and requirements, please see the KB article titled: LACP Support in Foundation.
If you are replicating data through DR (Data Replication page in Prism), then you have setup schedules for snapshots so they can be copied to the remote site on scheduled time. When you select a “snapshot” filed in one of the protection domains you created in the Prism UI, one of the fields is “Reclaimable space”. You may observe an spinning wheel continuously and the word “processing” on this filed for some or all snapshots, but you may also notice that snapshot(s) has already been taken and done. So why the spinning wheel for this filed? This field is lazy-calculated by Curator during full scans and populated afterward, so it takes sometime (may be few hours) to show up. Until Curator finishes calculating the value, the field shows Processing in Prism.
Hi, I have updated my Nutanix (AHV) license and checking license status, i see correct license expiry date applied to my clusters. I however still get alert from “Licese Standy Mode” details below Possible Cause The license file has not been applied after cluster summary file generation. Recommendation Apply a new license. When i try to reapply license by downloading csf file, uploading csf file i get error: You’ve uploaded an older/inactive cluster summary file(CSF). To continue, download the latest CSF from Prism Element or Prism Central Help please
Nutanix AOS offers simplicity in managing traditional complex infrastructure tasks. From Virtual machine management, Storage operations, replication - and of course Cluster software and hardware upgrades. As Infrastructure admins, we are well aware of the operational pain points, when it comes to upgrading: Hypervisor Upgrades Storage OS upgrades Firmware Upgrades Management software upgrades the list goes on… With Nutanix One-Click upgrades, customers can upgrade software components and hardware components easily. Software and Firmware needs to be downloaded from Nutanix repositories - which is why it is important to understand what Network Ports are required to be open or can be opened on demand to check for upgrades. Following KB from Nutanix Portal lists the required network ports for different services and upgrade repos endpoints: Recommendation on Firewall Ports Config
@Mutahir has already shared some insights on NCC checks in Keeping the Lights Green - NCC - Hardware Checks. Today I would like to bring up two important aspects of the tool. There may be a time where you receive an alert triggered by a regularly executed NCC check. Oftentimes the alert will have a reference to a KB article. You read the KB and it does not make any sense. Naturally, you raise a case with the Nutanix support team or commence the journey across vast space of the Internet in the search for an answer. The very first thing Nutanix support engineer will do is verify if the environment is running the latest version of NCC checks, and if it’s not, they will proceed with the NCC upgrade. More often then not, the alert will clear after the NCC upgrade. Why is it so? NCC is a powerful tool that is developed and maintained by a team of professionals. With their help the tool evolves and grows, more checks are introduced, issues are resolved and algorithms are improved. Thus it i
Below are the top knowledge base articles for the month of November 2019. KB 4141 - Alert - A1046 - PowerSupplyDown KB 4116 - Alert - A1187, A1188 - ECCErrorsLast1Day, ECCErrorsLast10Days KB 1540 - What to do when /home partition or /home/nutanix directory is full KB 7503 - G6, G7 platforms with BIOS 41.002 -DIMM Error handling and replacement policy KB 4409 - LCM: (LifeCycle Manager) Troubleshooting Guide KB 1113 - HDD/SSD Troubleshooting KB 4541 - Alert - A101055 - MetadataDiskMountedCheck KB 4158 - Alert - A1104 - PhysicalDiskBad KB 2090 - AHV | Host and Guest Networking KB 4519 - NCC Health Check: check_ntp KB 1888 - NCC Health Check: storage_container_mount_check KB 4188 - Alert - A1050, A1008 - IPMIError KB 1507 - Alert IPMI IP address on Controller VM was updated to ... without following the Nutanix IP Reconfiguration procedure, can be misleading KB 4273 - NCC Health Check: aged_third_party_backup_snapshot_check KB 3523 - How to create a Phoenix ISO or AHV ISO from a CVM or Foun
Below are new knowledge base articles published on the week of November 24-30, 2019. KB 8302 - Pre-Upgrade Check : test_is_hyperv_nos_upgrade_supported KB 8303 - Pre-Upgrade Check : test_if_cau_update_is_running KB 8499 - Security - Nutanix definitions for most common STIGs KB 8555 - Launching a blueprint by using the simple_launch API fails after Prism Central is upgraded to 5.11 KB 8616 - "Restore" screen under ASYNC DR is misaligned if entity have long name KB 8618 - PD: Trying to to 'deactivate-and-destroy-vms' operation got error 'Error: Unexpected application error kInvalidAction raised' KB 8619 - Genesis may not start with error 'Received multiple ips for interface bound to ExternalSwitch' KB 8621 - Alert - A400101 - NucalmServiceDown KB 8622 - Alert - A400102 - EpsilonServiceDown KB 8629 - Calm - Jenkins deployment is stuck at "Installing: ssh-credentials" and fails without error messages KB 8639 - AHV | Never-schedulable node CVMs are not shown in the VM in Prism. KB 8641 - De
Hi I have a concern with the data resilience in Nutanix Cluster about rebuild the data in 2 scenarios. When a node is broken or failure, then the data will be rebuilt at the first time, the node will be detached from the ring, and I can see some task about removing the node/disk from the cluster. The whole process will used about serveral minutes or half hour. It will last no long time to restore the data resilience of the cluster. When I want to remove a node from the cluster, the data will also be rebuilt to other nodes in the cluster. but the time will be last serveral hours or 1 day to restore the data resililence. Seems remove node will also rebuild some other data like curator,cassandra and so on. but Does it will last so long time, hom many data will be move additionaly ? and What the difference for the user data resilience for the cluster?
Hi all, I have some questions that I’m trying to answer but … ;) So if you can explain to me or point me to a part of some resources # Questions Is it recommended or mandatory to configure containers as ReplicationFactor-3 when the cluster is RedundancyFactor-3 In case of ReplicationFactor-3, when reading, how many checks are done to validate data correctness? In a RedundancyFactor-2 only 1 failure is tolerated, the cluster will still work with (e.g) 2 Zookeeper. In RedundancyFactor-3 there is 5 Zookeeper, so why we can’t tolerate up to 3 failure? What are the limitations for which it is not possible to migrate VMs between containers without the export/import method? How the cluster will behave in case of network separation issue (e.g. 4 nodes can communicate and 4 other too)? If I a have 2 Guest VM in the same Vlan, will they communicate through the OVS br0 or the traffic will go till the external switch and come back to the cluster? With the bond0 (br0.up) interface having 2 links
Nowadays everyone is concerned about the security of their infrastructure as they should be. The Nutanix document referenced below contains an overview of the security development life cycle (SecDL) and host of security features supported by Nutanix. It also demonstrates how Nutanix complies with security regulations to streamline infrastructure security management. In addition to this, this guide addresses the technical requirements that are site specific or compliance-standards (that should be adhered), which are not enabled by default. https://portal.nutanix.com/#/page/docs/details?targetId=Nutanix-Security-Guide-v510:Nutanix-Security-Guide-v510
We have had a couple of instances recently when making network changes that have affected our clusters. This caused a restart on the lead host due to it detecting a network loss and then resulted in system outages. The cluster is configured with dual networks ports in active and passive mode and the understanding was that it would switch if any change or failure was detecetd without producing error events and systems down.
All clusters will need to be upgraded at a point. If you have metro availability enabled in your environment, you will need to follow the best practices guide lines for it: https://portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v510:wc-metro-availability-upgrade-considerations-r.html Specifically Nutanix supports the following replications: Between N to N-2 major versions and vice versa for STS to STS versions or LTS to STS versions. Between N to N-1 major versions and vice versa for LTS to LTS versions.
Nutanix Pulse HD provides diagnostic system data to Nutanix support teams to deliver pro-active, context-aware support for Nutanix solutions. The Nutanix cluster automatically and unobtrusively collects this information with no effect on system performance. Pulse HD shares only basic system-level information necessary for monitoring the health and status of a Nutanix cluster. OK, this is all great but let’s get to real benefits, shall we? Q 1. Why would you enable it? A. Well, for several reasons: So that if you have an issue that you raise with Nutanix support we would be able to start looking at the data about your cluster immediately without making you answer many questions that are often crucially important for a prompt issue resolution. In essence, to reduce the amount of time it takes to fix the problem. So that when Nutanix engineer requires logs they would be able to collect them themselves while you would focus on what you need to do. No frustration with logs collection,
Hi all, did you guys know that AHV to ESXi conversion is only possible from Prism (Convert cluster option) if the cluster was earlier Esxi converted to AHV and being converted back to ESXi. AHV to ESXi conversion is not supported from the Prism. Some pre- requisites for converting AHV to ESXi via prism are-https://portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v55:man-cluster-conversion-requirements-limitations-r.html#nref_amd_hlp_k5 However one could always re-image an AHV node to ESXi manually. Before converting the cluster, customer should always migrate the running vms on the AHV cluster to their DR cluster. This can be done using Protection Domains. More information on Protection Domains can be found below: https://portal.nutanix.com/#/page/docs/details?targetId=Prism-Element-Data-Protection-Guide-v511:Prism-Element-Data-Protection-Guide-v511 or also refer KB-3059 https://portal.nutanix.com/#/page/kbs/details?targetId=kA03200000098T7CAI Once cluster is c
Hi all, I’m very new to Nutanix, and pretty new to Ansible. I’ve been tasked with updating / installing guest tools on any machines that need them, and they’d prefer to do it via Ansible. I’d like to be able to have Ansible use the uri module to grab the UUID of a given VM, or grab a list of UUID’s and the associated VM; however I’m having a lot of trouble parsing this information out in a way that Ansible can actually use it. Does anyone have experience with this? Or at least can tell me that there’s a better way to be doing this? Thanks!
Below are new knowledge base articles published on the week of November 17-23, 2019. KB 8193 - NCC Health Check: secure_boot_check KB 8223 - NCC Health Check: sed_key_availability_check and sw_encryption_key_availability_check KB 8513 - VSS Snapshot Fails for Windows VM having Dynamic Volume with Multiple Disk Extents KB 8514 - NCC Health Check: fs_inconsistency_check KB 8569 - Accessing Prism from a browser on MacOS 10.15 "Catalina" blocked by ERR_CERT_REVOKED error. KB 8571 - LCM firmware update unable to commence as VMs are unable to migrate off KB 8580 - When configuring a remote site configuration (Physical Cluster), "vStore Name Mapping" field will show an error "Remote site is currently not reachable. Please try again later." KB 8594 - How to change a snapshot's expiration time to indefinite KB 8600 - Calm License Changes and Grandfathering of Existing Users KB 8614 - Nutanix Files - Deleting TLDs in a NFS distributed share Note: You may need to log in to the Support Portal to v
Nutanix REST API gives flexibility to a developer or an administrator to create scripts which can execute administrative jobs on a Nutanix cluster. Using the API, you can request information about different entities in the cluster or even change some configuration. Everyone at the end of the day wants to make sure that their infrastructure is fully secure. So what about Authentication? What kind of authentication does REST API require? Nutanix support multiple authentication options. Want to know about them and how to configure and use them in your scripts? Give the following KB a readKB-2257 Want to know more about Nutanix REST API and different dev tools we provide? Log on to https://www.nutanix.dev/ Want to know more about Nutanix REST API Explorer and play around(with caution) with the APIs in your Nutanix Cluster? REST API Explorer
In this post we will go through the steps required to shut down and power off all hosts in a VMware vSphere cluster on AOS to perform maintenance or other tasks such as physical hardware relocation. Note: Upgrade to the most recent version of NCC before proceeding with the following steps. Points to consider before attempting to shutdown a vSphere Cluster running atop AOS:Cluster health and Resiliency (Prism Dashboard) vCenter hosted on the same AOS cluster or outside? Any on-going / running protection domain replication? User Virtual Machines shutdown CVM – Controller virtual Machines shutdown Access to ESXi Hosts (SSH & ESXi Host Client)Steps to ShutdownSSH to a Nutanix Controller VM (user : nutanix) and run an NCC health check prior to the scheduled shutdown. If there are any errors or failures, check the relevant KB or contact Nutanix Support.ncc health_checks run_allShut down all the VMs in the Nutanix cluster. Verify if vCenter VM is running on the same Nutanix cluster that y
Maintaining a big infrastructure requires planning and knowing your architecture inside and out.Let's say, sometimes you need to check the upgrade history of different components in your cluster.Sometimes you want to know when you last upgraded AOS or Nutanix Files.The timestamp of the upgrade.Let's say you want to plan a future upgrade and creating a time-line and want to know when you upgraded the software last and at which version.We all need this information, either to maintain a proper record of our cluster infrastructure or to plan the future upgrade.So how can we achieve this?Give the following Knowledge base article a read to find out the upgrade history of different components in your cluster and plan your future upgrades accordingly.KB-1151 Want to know how upgrades work in Nutanix? Try going through the following Knowledge Base to understand how upgrades work in Nutanix architecture.KB-6945 Have a question regarding upgrade workflow? Drop a comment and let’s start a conversa
How do I enable copy and paste in VMware for text in AHV Console?
Below are new knowledge base articles published on the week of November 10-16, 2019.KB 8349 - AHV | VM Goes Offline After Migration KB 8415 - Alert - A1132 - EntitiesSkippedDuringRestore KB 8512 - VM does not boot from Windows OS mounted as CD-ROM if VM already has OS installed on DISK Drives KB 8522 - LCM inventory failure "The requested URL returned error: 404 Not Found" in a dark site LCM environment KB 8525 - Alert - A130202 - AtlasNetworkingHostMemoryReservationFailure KB 8526 - Alert - A130201 - AtlasNetworkingHostConfigError KB 8527 - Pre-check:test_vcenter_connectivity KB 8540 - Pre-Upgrade Check: test_version_check KB 8544 - Pre-Upgrade Check: test_co_nodes_present KB 8551 - DiscoveryOS: Cluster Expansion - Cannot add or discover nodes : IPMI console displays phoenix prompt KB 8553 - How to change TimeZone in Move VM KB 8559 - Alert - A1182 - ShellvDisksAboveThresholdNote: You may need to log in to the Support Portal to view some of these articles.
Nutanix offers a distributed data and control plane, so it’s fairly easy to start, stop and graceful shutdown a cluster. Even, in abnormal / dirty shutdowns, Nutanix cluster has powerful self-healing capabilities - as all data & meta-data is distributed across the cluster which significantly reduces the chances for data corruption or data-loss. However, as a Nutanix Cluster hosts business critical data and applications, it is important to ensure all services stop in a graceful manner and all data + meta-data is consistent. This allows the cluster to be restarted in a healthy - usable state later.There can be several reasons to gracefully shutdown a running Nutanix AHV & AOS Cluster. When shutting down a Nutanix Cluster, following order needs to be followed:User VMs Shutdown AOS Shutdown - Data services / Cluster components CVM Shutdown Hypervisor ShutdownIn this post, we will focus on the shutdown process for a Nutanix Cluster running with AHV (Acropolis Hypervisor).Points to C
Hi,Our metro cluster is running AOS5.10.8.1 (since yesterday’s update from AOS 5.10.8) on x8 NX-8035-G6. During an LCM update for bios and HDD, a node failed due to a bad DIMM. The node pair were eventually evicted from the meta data ring at which point support were supplying a replacement DIMM. After a lot of coaxing both nodes were back in the cluster and hosting vSphere6.5 without any apparent issue, however, since the outage we are being warned in PC\PE and NCC that the active and standby PDs are not mounted on all nodes, even after manually ensuring that the containers are available and connected correctly. Support have advised that this is a known NCC error and just wondered if anyone else in the community has experienced this issue? Thanks in advance
Hi Team, Please let me know if you can share the Detailed Migration plan from ESXi to AHV. Please share the document or the link. Regards.
Let’s say you are managing an Infrastructure and have a range of networks defined in your environment and now you have to delete a few.You go to Prism, try deleting the network but get a generic error.So what’s going wrong here?Why can’t you delete a network?Well, one possibility is that if there are NIC connected to the network, it won’t let you delete the network, kind of like a guard-rail.So what can you do to mitigate it?Try giving the following KB a read KB-8234 Want to know more about AHV Networking? AHV Networking Best Practices Still confused about AHV Networking?Drop a comment and let’s start a discussion.
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.