Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
Let’s say you have connected a third-party storage device to a Nutanix Environment and now wondering whether it’s recommended or not. Nutanix Architecture is designed specially keeping in mind local storage settings and configuring any third-party storage devices can cause some serious issues in your environment. There are two different situations that might occur Permanent Device Loss All Paths Down Want to know the issues in detail and how to identify them Check out the KB-8674 In brief, connecting third party storage devices to a Nutanix architecture is not recommended unless and until especially advised by a Nutanix Team for certain scenarios.
Below are new knowledge base articles published on the week of December 1-7, 2019. KB 8546 - Pre-check: test_nsx_configuration_in_esx_deployments KB 8624 - PulseHD shows RED if dmidecode.exe is missing KB 8631 - Accessing Prism Via Citrix NetScaler (ADC) KB 8669 - Nutanix Files - long filename isn't supported yet. KB 8671 - How to determine which M.2 device failed on the node KB 8680 - Metro - Recovery procedure after two-node down scenarios Note: You may need to log in to the Support Portal to view some of these articles.
Let's say you have bought a new node(Yayy) and want to expand your cluster. Nutanix provides you with a GUI method to expand your cluster with ease and simplicity. Confused regarding the steps and the requirements before adding it to your cluster? Let's break down the queries to help you understand cluster expansion and the requirements in detail. When can you add the node in your cluster? If the node is physically present in the cluster network subnet, you can go ahead with cluster expansion. What should I check before cluster expansion? The cluster expansion process compares the AOS version on the existing and new nodes and performs any upgrades necessary for all nodes to have the same AOS version. What are the guidelines before proceeding with cluster expansion? What are the specific considerations I should keep in mind in case the hypervisor being AHV, Hyper-V or Esxi? Please go through the documentation to understand the guidelines in-depth. Cluster Expansion Guide Wh
Hi all, Local Replication is a process in which multiple copies of data are stored within a storage container. These copies exist for fault tolerance. Snapshots are placed locally on the same cluster as the source VM. Thus, If a physical disk fails, the cluster can recover data from another copy. The cluster manages the replicated data, and the copies are not visible to the user. So, what is the difference the Replication Factor option? Because RF is used too for fault tolerance in case of a physical disk failure (or node, ...) Thanks
Every once in a while due to network infra structure changes or because you have to physical move the cluster to another location, you may have to modify the Cluster IP. This includes CVM, Hypervisor and IPMI ip addresses, netmask and default gateways Unfortunately this operation requires taking some down time as you will need to stop the cluster for the duration of change. Before you start, you need to: 1- Clearing the external virtual ip address of the cluster , and setting new ip address for it 2- Ensuring that the Ntp and Dns servers of the cluster are reachable from new CVM ip address and if they are going to be different, remove the old addresses and add the new ones 3- Check that all hosts are part of metadata store You need to consider 3 different scenarios: 1- Change the IP addresses of the CVMs in the same subnet. 2- Change the IP addresses of the CVMs to a new or different subnet. 3- Change the IP addresses of the CVMs to a new or different subnet if you are moving the cl
Foundation is how we build and configure Nutanix clusters and many customers would prefer to make use of more advanced network technologies like LACP to improve the cluster performance and provide redundancy. LACP increases bandwidth, provides graceful degradation as failure occurs, and increases availability. It provides network redundancy by load-balancing traffic across all available links. If one of the links fails, the system automatically load-balances traffic across all remaining links. Foundation 4.2 introduces LACP support for the standalone Foundation. For more information about the supported hypervisors and requirements, please see the KB article titled: LACP Support in Foundation.
If you are replicating data through DR (Data Replication page in Prism), then you have setup schedules for snapshots so they can be copied to the remote site on scheduled time. When you select a “snapshot” filed in one of the protection domains you created in the Prism UI, one of the fields is “Reclaimable space”. You may observe an spinning wheel continuously and the word “processing” on this filed for some or all snapshots, but you may also notice that snapshot(s) has already been taken and done. So why the spinning wheel for this filed? This field is lazy-calculated by Curator during full scans and populated afterward, so it takes sometime (may be few hours) to show up. Until Curator finishes calculating the value, the field shows Processing in Prism.
Hi, We want to throttle replication bandwidth. In the Remote Site config in Data Protection, the option is to throttle B/W in MBps. Is this Mbps (megabits per second) or MB/s (Megabytes per second)? The unit implies megabytes per second (i.e. the "B" is capitalised), but that is normally a unit of throughput not bandwidth... Obviously there is an order of magnitude in difference so its important we interpret this correctly!
Hello, Lately I have been adding the Nutanix SCOM Management Pack version 2.4.0.0. From the start on it worked fine as I handed over the Cluster information via Nutanix Cluster Discovery. But now, a few weeks later, SCOM will not display any performance data of the clusters anymore. I am not able to find out where exactely caused this problem. I have different Clusters AHV and ESXI as well as multiple OSVersions running. Currently SCOM displays only one Cluster with it's information. From the other Clusters, there is no performance data shown. All other clusters were discovered correctely (and the same way) but won't show their data anymore. My situation looks like this - this dashboard is showing the clusters information. All other clusters (and their nodes) are not showing data - the dashboard stays empty. What could have happend that the data is not reaching SCOM or SCOM not displaying it anymore? Could it be too much Clusters on my system? I have 13 Clusters with (together)
Hi I have a concern with the data resilience in Nutanix Cluster about rebuild the data in 2 scenarios. When a node is broken or failure, then the data will be rebuilt at the first time, the node will be detached from the ring, and I can see some task about removing the node/disk from the cluster. The whole process will used about serveral minutes or half hour. It will last no long time to restore the data resilience of the cluster. When I want to remove a node from the cluster, the data will also be rebuilt to other nodes in the cluster. but the time will be last serveral hours or 1 day to restore the data resililence. Seems remove node will also rebuild some other data like curator,cassandra and so on. but Does it will last so long time, hom many data will be move additionaly ? and What the difference for the user data resilience for the cluster?
Nutanix AOS offers simplicity in managing traditional complex infrastructure tasks. From Virtual machine management, Storage operations, replication - and of course Cluster software and hardware upgrades. As Infrastructure admins, we are well aware of the operational pain points, when it comes to upgrading: Hypervisor Upgrades Storage OS upgrades Firmware Upgrades Management software upgrades the list goes on… With Nutanix One-Click upgrades, customers can upgrade software components and hardware components easily. Software and Firmware needs to be downloaded from Nutanix repositories - which is why it is important to understand what Network Ports are required to be open or can be opened on demand to check for upgrades. Following KB from Nutanix Portal lists the required network ports for different services and upgrade repos endpoints: Recommendation on Firewall Ports Config
I am looking for a something that i can setup in an automated task on a server to poll for any active replications for Protection Domains and if true to pull information and email it. The output i am looking for in the email would be something like below. Protection Domain : ProtectionDomainName Replication Operation : Sending Start Time : 03/11/2019 12:00:02 EDT Remote Site : RemoteSiteName Snapshot Id : 2918635 Bytes Completed : 444.08 MiB (465,653,447 bytes) Snapshot Size : 2.57 GiB (2,760,598,528 bytes) Complete Percent : 95.38689 If anyone already has something like this setup that would be awesome, my scripting skills are slim to none so any help would be awesome.
Hi all, I’m very new to Nutanix, and pretty new to Ansible. I’ve been tasked with updating / installing guest tools on any machines that need them, and they’d prefer to do it via Ansible. I’d like to be able to have Ansible use the uri module to grab the UUID of a given VM, or grab a list of UUID’s and the associated VM; however I’m having a lot of trouble parsing this information out in a way that Ansible can actually use it. Does anyone have experience with this? Or at least can tell me that there’s a better way to be doing this? Thanks!
I'm running Prism Central 5.7.1.1 and have recently upgraded 6 of our clusters to 5.5.7.1. Of those 6 clusters, 3 of them are still showing an Upgrade Status of 'Upgrading' in Prism Central. It's been a over a month for one cluster and I've restarted Prism Central appliance to no avail - has anyone else seen this?
I've seen servers that have active and inactive in replication . when I check the replication settings. there are inactive and active vm's in job. so i didnt understand that what is mean inactive vm's. can i delete it . if there is no replication in the other region, they take up storage space and if I delete them I can save space from storage space Also i have another question is on vcenter 6.5. i see free size is 4 tb but on nutanix mangement screen free size is 11 tb. why do i see it so different. If the virtual machine is deleted, there is a space recovered on the storage side.
Hey everyone, I am required to upgrade from ESX 6.7 U1 to U2 to resolve a bug that prevents me from going back more than 8-10mins to look at VM performance, however, Nutanix has only just now tested U3. My question is: Has anyone migrated to U3 yet and if so how has it been so far?
Need help to change the IP address , Subnet mask and gateway of CVM, IPMI, ESXi Hosts and vCenter AOS 5.0.4 ESXi 5.5.0 Build number - 1331820 ncc 3.1.2 vcenter 5.5.0 Build number - 2442329 total number of nodes : 5 1) what is the Order to change the IP address, subnet mask and gateway of 5 CVM, 5 ESXi Hosts, 5 - IPMI, 1 - vCenter 2) do i need to stop the cluster for CVM - IP, subnet mask and gateway change? 3) do i need to stop the cluster for ESxi Hosts - IP, subnet mask and gateway change? 4) is there any workaround to update the CVM - IP address, subnet mask and gateway without stopping the cluster 5) reason for stopping the cluster while changing the IP address, subnet mask, gateway 6) does this AOS version will support CVM and ESXi IP change 7) what has to be check before and after IP address, Subnet mask, gateway change of CVM, ESXi, IPMI and vCenter 8) What is the Procedure of changing the CVM IP address, Subnet mask and gateway (AOS 5.0.4/Esxi 5.5.0) 9) What is the Procedur
@Mutahir has already shared some insights on NCC checks in Keeping the Lights Green - NCC - Hardware Checks. Today I would like to bring up two important aspects of the tool. There may be a time where you receive an alert triggered by a regularly executed NCC check. Oftentimes the alert will have a reference to a KB article. You read the KB and it does not make any sense. Naturally, you raise a case with the Nutanix support team or commence the journey across vast space of the Internet in the search for an answer. The very first thing Nutanix support engineer will do is verify if the environment is running the latest version of NCC checks, and if it’s not, they will proceed with the NCC upgrade. More often then not, the alert will clear after the NCC upgrade. Why is it so? NCC is a powerful tool that is developed and maintained by a team of professionals. With their help the tool evolves and grows, more checks are introduced, issues are resolved and algorithms are improved. Thus it i
We are using a Prism Central Installation with several Clusters in different Sites. For the Administrators on Site, we want to configure a user on the site based prism elements instance, which is a viewing User, with the following additional permissions on the local cluster / Prism Elements: Viewing, Power on and off for all virtual Machines and also a one click shutdown of the whole cluster in case of power outages. We don't want to use the Prism Central User Configuration for following reasons: Prism Central is not accessible in case of line outages or power failures in our sites. And for security reasons and limited skills of the staff on site, I really appreciate not to give admin permission to the site staff. Any Idea, how to deal with this situation ? Thanks Oliver
Hi Team, We have multiple Nutanix clusters with Dell XC servers installed with ESXI servers. We found an user account "PTAdmin" present in Dell iDRAC servers. Is there anyway to disable the account from CVM or Hypervisor since we have 100+ nodes in the infra. Thanks.
Is there a minimum or recommended switch port buffer size? I plan on using eth0 and eth 1 to be connected to two different 10G switches but was wondering if there is a minimum buffer size or switch port requirement?
Below are the top knowledge base articles for the month of November 2019. KB 4141 - Alert - A1046 - PowerSupplyDown KB 4116 - Alert - A1187, A1188 - ECCErrorsLast1Day, ECCErrorsLast10Days KB 1540 - What to do when /home partition or /home/nutanix directory is full KB 7503 - G6, G7 platforms with BIOS 41.002 -DIMM Error handling and replacement policy KB 4409 - LCM: (LifeCycle Manager) Troubleshooting Guide KB 1113 - HDD/SSD Troubleshooting KB 4541 - Alert - A101055 - MetadataDiskMountedCheck KB 4158 - Alert - A1104 - PhysicalDiskBad KB 2090 - AHV | Host and Guest Networking KB 4519 - NCC Health Check: check_ntp KB 1888 - NCC Health Check: storage_container_mount_check KB 4188 - Alert - A1050, A1008 - IPMIError KB 1507 - Alert IPMI IP address on Controller VM was updated to ... without following the Nutanix IP Reconfiguration procedure, can be misleading KB 4273 - NCC Health Check: aged_third_party_backup_snapshot_check KB 3523 - How to create a Phoenix ISO or AHV ISO from a CVM or Foun
Below are new knowledge base articles published on the week of November 24-30, 2019. KB 8302 - Pre-Upgrade Check : test_is_hyperv_nos_upgrade_supported KB 8303 - Pre-Upgrade Check : test_if_cau_update_is_running KB 8499 - Security - Nutanix definitions for most common STIGs KB 8555 - Launching a blueprint by using the simple_launch API fails after Prism Central is upgraded to 5.11 KB 8616 - "Restore" screen under ASYNC DR is misaligned if entity have long name KB 8618 - PD: Trying to to 'deactivate-and-destroy-vms' operation got error 'Error: Unexpected application error kInvalidAction raised' KB 8619 - Genesis may not start with error 'Received multiple ips for interface bound to ExternalSwitch' KB 8621 - Alert - A400101 - NucalmServiceDown KB 8622 - Alert - A400102 - EpsilonServiceDown KB 8629 - Calm - Jenkins deployment is stuck at "Installing: ssh-credentials" and fails without error messages KB 8639 - AHV | Never-schedulable node CVMs are not shown in the VM in Prism. KB 8641 - De
Hi all, I have some questions that I’m trying to answer but … ;) So if you can explain to me or point me to a part of some resources # Questions Is it recommended or mandatory to configure containers as ReplicationFactor-3 when the cluster is RedundancyFactor-3 In case of ReplicationFactor-3, when reading, how many checks are done to validate data correctness? In a RedundancyFactor-2 only 1 failure is tolerated, the cluster will still work with (e.g) 2 Zookeeper. In RedundancyFactor-3 there is 5 Zookeeper, so why we can’t tolerate up to 3 failure? What are the limitations for which it is not possible to migrate VMs between containers without the export/import method? How the cluster will behave in case of network separation issue (e.g. 4 nodes can communicate and 4 other too)? If I a have 2 Guest VM in the same Vlan, will they communicate through the OVS br0 or the traffic will go till the external switch and come back to the cluster? With the bond0 (br0.up) interface having 2 links
We have had a couple of instances recently when making network changes that have affected our clusters. This caused a restart on the lead host due to it detecting a network loss and then resulted in system outages. The cluster is configured with dual networks ports in active and passive mode and the understanding was that it would switch if any change or failure was detecetd without producing error events and systems down.
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.