Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
Below are new knowledge base articles published on the week of June 7-13, 2020. KB 9456 - Alert - A400114 - PolicyEngineServiceDown KB 9468 - Different Docker hosts can see volumes created with Nutanix DVP KB 9478 - How to clear stuck LCM inventory tasks ? KB 9480 - Windows 10, version 2004 or Windows Server, version 2004 VMs may fail to boot on AHV KB 9487 - AHV | VM update operations initiated from Prism Central may fail with "Entity CAS version mismatch" error due to missing machine_type attribute KB 9492 - Move VM IP address may change post deployment KB 9494 - Disks from ISCSI volumes change drive letters on Windows VMs after upgrade AHV 2016* to 2017* Note: You may need to log in to the Support Portal to view some of these articles.
Hi all, A newbie question. It seems I still have an old crash dump directory on one of my AHV hosts. A ls -lahtr /var/crash shows the single directory from back in April. At the time the issue was resolved, and the faulty DIMM that was causing the issue was replaced. That said, clearly the dump file was not removed. In terms of cleaning this up, is it OK to delete the dump directory within the /var/crash directory and then rerun the ncc health check. Or is there a better method for clearing crash dumps from Nutanix clusters. Many thanks, Rob
Please answer the below questions : When one cvm goes down(in a 10 node cluster with RF3) for 20 mints, then Guest VM’s new write and read IO would be served by anther's CVM and all these will be traveling across 10 g network. A : Will those new IOs served from local copy via another CVM or IOs will be served from the replica copy B: Will cluster starts to build a new replica to accommodate the missing copy or RF3 Ques 2 - When new Write IO request comes, first it will write the data on Oplog then synchronously send to other CVM’s Oplog. All host in clusters having 2 SSD and 6 HDDs. Will Write IO process by both SSD in every host or only one SSD is hsving Oplog partition active at the same time ? I think there is only one oplog per CVM/host whether the host is aving all flash drive or 2 SSD and remaining HDD. But I am not sure about this statement, Please clarify..
Is it possible to recover data from an overwritten entity - does the overwritten VM or its storage still exist anywhere? I'm asking as I recovered a VM from a backup but realised in hindsight that the logs from the VM in its failed state would be useful for diagnosing the issue. Thanks in advance.
What is the DIMM error? A memory error is an event that leads to the logical state of one or multiple bits being read differently from how they were last written. For example, If 1 was written in a memory cell and while reading the same memory cell, it returns 0. Memory errors can be classified into two types: Soft errors, which randomly corrupt bits but do not leave physical damage. Soft errors are transient in nature and are not repeatable. Soft errors can be because of electrical or magnetic interference (e.g. due to cosmic rays, alpha particles, leakage, random noise). Hard errors, which corrupt bits in a repeatable manner because of a physical/hardware defect or an environmental problem. Hard error can also occur if DIMM is not seated properly. All memory systems in use in servers today are protected by error detection and correction codes. These server machines employ error correcting codes (ECC), which allows the detection and correction of one or m
First, let us understand what NTP (Network time protocol) is. An NTP server is a time server that is used to keep/sync the time in your cluster. An NTP server can be public or private depending on the strictness of your environment. To know how to configure NTP in your Nutanix cluster, take a look at- https://support-portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v5_16:wc-system-ntp-servers-wc-t.html After the NTP server is configured, the genesis leader becomes the NTP leader, which means that the genesis leader is syncing time to the NTP server and other CVMs are syncing time with the genesis leader. How NTP works in AHV:- It’s as simple as it gets. The AHV hypervisor takes the same server configured on the cluster and syncs the time with it individually. There are no extra steps required to configure the NTP server on the AHV hosts. How NTP works in ESXi:- The ESXi cluster does not take the server configured on the Nutanix cluster and it needs to be
Hello, I’m working with Rest API, to “automate” some operations with Ansible, trought URI Module and Jinja2. I need to create a Project, and set permission to a specific user. How can i do this with API V3? I can add user to a Project, but i can’t set the Role.. When i check trought the Web Interface, i see the user without role. Thanks!
Below are new knowledge base articles published on the week of May 31-June 6, 2020. KB 9442 - LCM BIOS/BMC Upgrade fails when node does not respond to IPMI power reset KB 9460 - Move: ESXi-AHV migration connection limits KB 9463 - Pre-upgrade check: Hypervisor Upgrade (test_host_upgrade_versions_compatible) KB 9467 - How to enable Karbonctl in Karbon darksite environment Note: You may need to log in to the Support Portal to view some of these articles.
Hi, everyone I’m trying to use LCM to update firmware of host machines on my cluster. It all went well until the post action phase. Error message said: ‘Operation failed. Reason: LCM failed performing action reboot_from_phoenix in phase PostActions on ip address xx’. I searched the KB and found KB9177,but it’s about ‘Mixed Hypervisor cluster‘, my cluster uses solely AHV so it doesn’t applied. Anyway I still tried to follow the KB9177’s suggestion and upgraded my cluster’s foundation to 4.5.3 and retried the LCM firmware update process on another host but still got the same ‘ LCM failed performing action reboot_from_phoenix’ error. I used the workaround provided in that KB to make the affected two hosts’ CVM out of maintenance mode. It works and the cluster is back to normal. Then I logged on to the affected hosts’ IMM and found out that actually the primary IMM2 firmware has already been updated by the LCM( the backup IMM2 firmware is not upgraded), and when I go to the LCM section
Hello, for LCM updates I see that pre-checks require DRS in fully automated mode. It’s possible to perform LCM updates without DRS in fully automated mode (in the event that a customer does not have DRS in his licenses) ? Thanks Manuel
Hello. I faced with limitation of number VMs in API reply. There is option “length” which can be added to request, but max value is 500 (I got it into experiments). We have around 1000 VMs in our environment. Is there solution?
Hello, We have a cluster with 4 nodes and I will perform the update from 5.9.4.2 to 5.10.10.1. a some minutes ago, I see that my cluster have a disk space problem and the resilience board is red with the “not resilience” information. In this case I will have a problem when perform the AOS upgrade? I not have time to resolve the space issue now… Please, send me any information.
is there any way to created additional users in nutanix move? -jai
I’m trying to setup a new bridge (br1). After ssh to the Nutanix AHV, I continue with input “ssh root@xxx.xxx.xxx.xxx” to get to the CVM, it ask me for a password. Is there a way to change/ continue with a different logon name as well at that? The password I input is being denied due a different username for the CVM.
Hello, comunity I’m been trying to figure out why the configuration with RF2 & N+1 reserves more capacity than just RF2. I understand that with RF2 all the data is duplicated along the cluster and the failure of one node is tolerated. In my scenario I have a 3-node cluster with 34TiB of effective capacity, so with RF2, I would have only 17TiB, everything is clear until that point. But the extent store assuming RF2 & N+1 gives me 11.3TiB, which is 1/3 of my available capacity. So...the question is: If I reserve capacity, by tripling all the data, that is for tolerate 1 additional node failure besides the 1st one (tolerated by RF2)? A silly conclussion would be that with RF2 & N+1 the cluster is able to tolerate 2 nodes and continue operating just with one node, but I know that is not posible. So, why assuming RF2 & N+1 reserves more TiB’s than just RF2? Please, I would appreciate the help. Thanks in advanced!
I have a W2k19 VM I created as a patched image. Idea is to patch it monthly. Shut it down and clone it. Then sysprep the clone and copy the disk to the Image Service. Then I can create VM’s off of it. I will repeat the colne after patching the ~master image monthly in order to keep a patched disk ready to spin up VM’s from. We are using AHV, version 5.10. Mt question is the “CLONE” a full, independant copy of the VM ? TIA -S
Hi community, I’m fairly new to interacting with Nutanix for automation and was hoping my VM build task would be a simple process but I’ve hit a wall with the guest tools install. I’m building a RDS environment where I need as little human interaction as I can get to make administration easier so I’m building my host VM using SCCM. I’ve multiple good reasons for approach but its mainly in support for application deployment via happening at build time. I’ve tried to install the guest tools manually but I believe they are failing due to the files being copied from a mounted ISO which has some dynamically created certificates involved so this isn’t going to work. I can see the ISO can be mounted using the API calls but every reference I’ve seen looks to need a username/password passing in the script and I’m not overly keen in hard coding creds into script. Most of what i’ve seen is a few years old so I wondered if there has been any changes to the approach to mounting guest tools? It
Below are the top knowledge base articles for the month of May 2020. KB 7503 - G6, G7 platforms - DIMM Error handling and replacement policy KB 4116 - Alert - A1187, A1188 - ECCErrorsLast1Day, ECCErrorsLast10Days KB 1540 - What to do when /home partition or /home/nutanix directory is full KB 7604 - Disk space usage for root on Controller VM has exceeded 80% KB 4141 - Alert - A1046 - PowerSupplyDown KB 4409 - LCM: (LifeCycle Manager) Troubleshooting Guide KB 4158 - Alert - A1104 - PhysicalDiskBad KB 1113 - HDD/SSD Troubleshooting KB 2090 - AHV | Host and Guest Networking KB 2473 - NCC Health Check: cvm_memory_usage_check KB 8792 - NCC checks: same_hypervisor_version_check, duplicate_cvm_ip_check, same_timezone_check, esx_sioc_status_check, power_supply_check, orphan_vm_snapshot_check giving ERR KB 4519 - NCC Health Check: check_ntp KB 2486 - NCC Health Check: cvm_mtu_check KB 4273 - NCC Health Check: aged_third_party_backup_snapshot_check KB 3357 - NCC Health Check: ipmi_
Below are new knowledge base articles published on the week of May 24-30, 2020. KB 9390 - Alert - A130338 - ServiceBadScore KB 9431 - LCM operation on ESXi hosts fails with "LAG is configured with X uplinks however there are Y NICs added in the teaming policy" Note: You may need to log in to the Support Portal to view some of these articles.
Giving thought to change your replication factor from 2 to 3? What are the impacts and things to consider? First, let’s take a look at what replication factor is. Redundancy factor is a configurable option that allows a Nutanix cluster to withstand the failure of nodes or drives in different blocks. By default, Nutanix clusters have redundancy factor 2, which means they can tolerate the failure of a single node or drive. So RF3 means cluster can tolerate the failure of 2 nodes or drive… Basic Maths isn’t it? Redundancy factor 3 has the following requirements: Redundancy factor 3 can be enabled at the time of cluster creation or after creation too. A cluster must have at least five nodes for redundancy factor 3 to be enabled. For guest VMs to tolerate the simultaneous failure of two nodes or drives in different blocks, the data must be stored on containers with replication factor 3. Controller VMs must be configured with a minimum of 28 GB(20 GB default+8 GB for the featur
New with Nutanix Move. Deplyed Nutanix Move and need to set IP up. Selecting new Nutanix-Move VM > Console » move login is required, which was not asked to set during MOVE setup. What log-on info needed to be set and where, or is there a standard logon there?
Dear all We run 2 Clusters with each 3 ESXi and 2 AHV Storage-only Nodes successfully. Now I want to get rid of VMWare, but all vms were built with vmware, of course I see them in prism. Did anyone ever just deinstall ESXi and replaced it with AHV in a running environment? Alternatively I could move all urgent vms to one Cluster and than install Nutanix from scratch and afterwards migrate the vms to the Nutanix only Cluster. thanks and best regards, Claudia
Below are new knowledge base articles published on the week of May 17-23, 2020. KB 4781 - A Windows 7 or Windows 10 User VM Fails to Power On with an NVIDIA vGPU Profile in ESXi 6.5b KB 9218 - Prism Element alternative UPN login fails KB 9357 - Oops - Server error when updating Container settings in Prism KB 9373 - AHV | 10Gbps NIC shows as 1Gbps despite auto negotiation enabled on Intel X550 NIC cards KB 9394 - Objects - Deployment might fail with Asymmetric routing / Policy Based routing configured in environment | Deployed cluster greyed out | IAM service is not healthy KB 9400 - AHV | Troubleshooting Virtual Machine boot failures KB 9403 - LEAP - Entity Sync tasks generated continously after PC upgrade to 5.16.1 KB 9407 - [Objects] Objects 2.1 Darksite Deployment failing at deploying metadata store KB 9420 - /tmp filled up by /tmp/paramiko_logs file KB 9422 - Error updating the current checks parameter. Health check schedule interval is invalid. KB 9425 - Official Guidan
Below are new knowledge base articles published on the week of May 10-16, 2020. KB 9304 - Object Store Deployment Failure: type[CREATE]:code: 400, message: Primary MSP deployment in failed state status_code: 400 KB 9329 - ESXi | 1 click upgrade failing with 'Could not find a trusted signer' KB 9333 - Prism Central - Adding same email address under both "report settings" as well as "schedule settings" causes duplicate emails to be sent to recipient KB 9349 - [Nutanix Objects][Common Deployment Failure Scenario] Failed to create Envoy VM | Prism Central is not able to reach MSP and Envoy VMs or slow image download KB 9354 - Error: Node XXX cannot be removed: Cluster needs at least 5 usable nodes KB 9368 - Karbon UI stuck at loading in case of old browser version KB 9377 - [Karbon] Darksite Deployment fails in Proxy Environment in ETCD Deployment Stage KB 9383 - AHV | manage_ovs tool shows error "Failed to send RPC request. Retrying." when trying to change bridge configuration KB 9389 - D
Can any one list out how many types of Foundation are there eg Foundation Pro, bare metal ..etc.
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.