Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
Hi AllWhen I try to install AOS 6.5.2 I get the following errors. IOError: [Errno 2] No such file or directory: '/etc/nutanix/factory_config.json' 2023-03-07 13:19:54,652Z CRITICAL svm_rescue:926 No suitable SVM boot disk found. 2023-03-07 13:19:54,652Z INFO svm_rescue:114 exec_cmd: sync; sync; sync 2023-03-07 13:19:54,658Z INFO svm_rescue:114 exec_cmd: umount -R /mnt/disk 2023-03-07 13:19:54,663Z INFO svm_rescue:114 exec_cmd: umount -R /mnt/data] 2023-03-07 13:19:52,159Z INFO Imaging thread 'svm' failed with reason [None] 2023-03-07 13:19:52,164Z CRITICAL Imaging thread 'svm' failed with reason [None] 2023-03-07 13:19:52,200Z ERROR Exception in running <InstallHypervisorKVM(<NodeConfig(172.16.150.9) @b5d0>) @ee10> Traceback (most recent call last): File "foundation/imaging_step.py", line 161, in _run File "foundation/imaging_step_hypervisor.py", line 47, in run File "foundation/imaging_step.py", line 353, in wait_for_event StandardError: Received "fatal" in waiting for ev
Below are new knowledge base articles published on the week of March 5-11, 2023.KB 12446 - NCC Health Check: pc_to_ahv_secondary_ip_reachability_check KB 14162 - Foundation Failing Due to Incorrect Date or TimeNote: You may need to log in to the Support Portal to view some of these articles.
My 3 node cluster consiting of 1065-G5 nodes is coming EOL this year. I will be replacing the hardware with a NX-3360N-G8 cluster. What is the recommended workflow for replacing the hardware? One thing to note is the current cluster is licensed with Prism Pro and the new one will be licensed with Starter as I’m not using all the features in Prism Pro.Is the replacement as easy as expanding the cluster one node at a time, migrating the workloads and then removing the old cluster? Will there be any hiccups with difference in licensing?
Hey guysI am reinstalling three Lenovo HX5521 nodes and I got stuck on the below errorFoundation IP not set. Try running the “set_foundation_ip_address” script on the desktopI'm running the foundation vm on the same subnet of the IPMI interfaces connected directly to an unmanaged switch, without VLANs or any other configuration. connectivity is perfect. I already reviewed all the settings and tried to reimage foundation using ESXi and AHV. Both result in the same error.Im running;Foundation_VM-5.2.2 AOS euphrates-5.20.3 LTS VMware-ESXi-7.0.1 or AHV-20201105.2244I've been trying to solve this problem for three days now, but so far I haven't found any clues. Any help will be greatly appreciated. Thanks! 👷🏽
Below are new knowledge base articles published on the week of February 26-March 4, 2023.KB 13178 - NCC Health Check: cluster_node_count KB 13481 - Increasing /dev/sda3 filesystem partition size on Policy VM upgraded to 3.6.1 KB 13653 - Nutanix Database Service | Pulse Telemetry KB 14018 - Alert - A130365 - PauseStretchTriggeredByWitness KB 14032 - Upgrading to NDB 2.5.1 for HA-enabled setups fails KB 14303 - Unable to delete a Nutanix Storage Container KB 14326 - AOS upgrade to 6.1.x (or above) stuck on network segmentation enabled cluster with mixed hypervisors. KB 14366 - Nutanix Database Service - MongoDB brownfield to greenfield conversion fails after upgrade from a version lower than 2.5.1 to a higher one KB 14367 - IPMI Network Configuration Requirements & Best Practices KB 14372 - Data Lens - how to re-enable file server after been disabled KB 14373 - Storage Container Usage may increase after AOS upgrade in ESXi clusters with Thick Provisioned virtual disks larger than 4Ti
Hello,We have an unsupported configuration on 2016 stretch cluster: 3 NIC teaming with LACP → should be 1 NIC Teaming with a switch independent The goal is to upgrade to server 2022, we have the manual option from TAC to upgrade to 2019 than 2022.Note based on what we have seen we have the following points:-From the doc. https://portal.nutanix.com/page/documents/details?targetId=Acropolis-Upgrade-Guide-v6_6:upg-cluster-upgrade-recommend-hyperv-r.html Upgrade to Windows Server 2022 Hyper-V from an LACP enabled Hyper-V 2019 cluster is not supported. Enabling Link Aggregation Control Protocol (LACP) for your cluster deployment is supported when upgrading hypervisor hosts from Windows Server 2016 to 2019. On the other hand in doc. https://portal.nutanix.com/page/documents/kbs/details?targetId=kA00e000000LLI8CAO the LACP is not supported on 2016 and 2022, since we have seen the following error as seen below (the ISO image was 2019 for the upgrade, not 2022 we can not upgrade to 2022 direc
Using phoenix to reimage a node after a failed satadom, after phoenix is loaded, I uploaded AHV ISO so it can proceed to install AHV/AOS. It turns out I uploaded the lcm_ahv_el7.nutanix.20201105.2267.tar.gz instead of the AHV-DVD-x86_64-el7.nutanix.20201105.2267.iso. The install is now hung at 75% Node discovery succeeded. I have tried to install the proper iso, but it immediately failed at that same progress. How can I cancel this install so I can try again?I did find this error in the foundation log. It looks like the error happens if you use the html 5 client to mount the phoenix iso, but I am using the java client.
I am able to create a Move plan using the API and Ansible for migrating a VM from vSphere to Nutanix. However, I am unable to specify the guest credentials for the guest operations (like removing VMware Tools from the replicated vm during the migration) in the actual rest call (see code below) as is possible using the GUI. I cannot find any way in the API documentation on how to get the credentials into the request. Hope somebody can assist.- name: "Create Move Plan: Task 1.2a - Create Move Plan for OTA" uri: url: "https://{{ move_server_ota }}/move/{{ mapi_version }}/plans" method: POST validate_certs: no force_basic_auth: yes headers: Authorization: "{{ move_token_ota }}" body_format: json body: | { "Spec": { "Name": "{{ move_host_name | lower }}", "NetworkMappings": [ { "SourceNetworkID": "{{ move_source_network_id[0] }}", "TargetNetworkID": "{{ move_ota_target_network_uuid}}",
We have our Nutanix 5 node cluster installed with VMware vSphere 7.0U3. RF=2 so I should be bale to withstand a 1 node failure. I would like to test the HA capability of VMware/Nutanix by taking a node offline abruptly and ensure everything works as expected before placing production workloads on it.I was thinking of placing a test VM on a node and then pulling the power cables and ensuring that VM is powered up on a remaining host.Does anyone have any better or preferred way to test HA?
Below are the top knowledge base articles for the month of February 2023.KB 1540 - [AOS Only] What to do when /home partition or /home/nutanix directory on a Controller VM (CVM) is full KB 7503 - NX Hardware [Memory] - DIMM Error handling and replacement policy KB 8885 - Alert - A15039 - IPMI SEL UECC Check KB 1381 - NCC Health Check: host_nic_error_check KB 4409 - LCM: Life Cycle Manager Troubleshooting Guide KB 3827 - Alert - A130087 - Node Degraded KB 1113 - HDD or SSD disk troubleshooting KB 2090 - AHV host networking KB 2473 - NCC Health Check: cvm_memory_usage_check KB 13870 - Prism Virtual IP is configured but unreachable alert and VIP becomes permanently unreachable observed on AOS 6.5.1 and later KB 3786 - Alert - A1081 - CuratorScanFailure KB 8514 - NCC Health Check: fs_inconsistency_check KB 4519 - NCC Health Check: check_ntp KB 13150 - NCC Health Check: cfs_fatal_check KB 5228 - NCC Health Check: pcvm_disk_usage_check KB 8094 - NCC Health Check: disk_status_check KB 4158 -
Below are new knowledge base articles published on the week of February 19-25, 2023.KB 13146 - Flow Network Security may block traffic using virtual IP address or forwarded traffic KB 13587 - List of IAM Migration Alerts during Upgrade to Prism Central pc.2022.9 KB 14017 - Alert - A130354 - FailoverTriggeredByWitness KB 14308 - Network performance of VXLAN interfaces inside VMs severely degraded if VM is running on AHV host with Broadcom BCM57414 NICs with Hardware Generic Receive Offload (GRO) feature enabled KB 14332 - Metric IO stats missing from Prism or the Insights Portal KB 14343 - Prism Central - Templates view is empty for LDAP Users on Prism Central with enabled Microservices Platform (CMSP) KB 14344 - Prism Central: report is missing values or shows only a dash (-) KB 14351 - How to obtain quota limits via API after enabling Calm Policy Engine KB 14353 - Unable to change the VPN route priority in Nutanix DRaaS KB 14369 - AHV Metro/SyncRep : Services on Pacemaker unable to ru
Hello, i have nutanix 3060g8 i have error:Committed memory update intent that is stuck on 50% i tried to reboot cvm and all nodes but i still have same problem.cvm memory are 64gb i tired:ecli task.list include_completed=noTask UUID Parent Task UUID Component Sequence-id Type Statusf0ffcb92-96b1-426f-51d7-204729b269fe kGenesis 1 kCvmreconfig kRunning progress_monitor_cli --entity_id="f0ffcb92-96b1-426f-51d7-204729b269fe" --deleteany suggest ?
Hi Team,Which Firewall ports will be required for new node we are adding node in the existing cluster.There is VMware cluster with Nutanix cluster.kindly guide and suggest.
Hello eveyone,following my last post I eventually got my C node to be alive. (it was not booting nor being visible whatever I tries, always down).I can ping it with both IPv4 and IPv6 addresses, connect to it in SSH, run commands etc.Now my situation is : while I was struggling with my node down, I removed it from the cluster with Prism. I think I could bring it back after… But it does not work.As you can see here, the C node (#3) is now missing :I use the “Expand cluster” tools in Prism Element, I add the node manually, it is detectedI check the mode, validate and it starts to expand.But after a few minutes I get the error :Failure in pre expand-cluster tests. Errors: Failed to get HCI node info using discovery It seems like the cluster refuse to consider the node as “free”, or the node itself refuse to join because it thinks that it is still in the cluster.Thank you very much for any help you could provide
I got existing 3 nodes with Dell XC740 running with AHV. And I’ve order new 3 nodes x Dell XC750 but it came with factory install ESXi 7.0.Question is any special instruction do before adding new node to the existing cluster and I would like to new node XC750 to run the same AHV hypervisor in existing cluster too. So, can I follow expand step here Prism 6.5 - Expanding a Cluster (nutanix.com)?Thank you.
Below are new knowledge base articles published on the week of February 12-18, 2023.KB 14236 - PSOD on Ice Lake processor platforms after upgrading to ESXi 7.0 U3i due to an issue with microcode 0x0d000375 KB 14297 - Security Dashboard throws the error "Enable microservices infrastructure with internet connectivity for Security Dashboard to work" KB 14309 - NDB - MySQL software profile creation fails for commercial MySQL versions KB 14310 - NDB - Failed to create the software profile with error message: "local variable \'db_version\' referenced before assignment" KB 14318 - User defined Life cycle rule does not work in object bucketsNote: You may need to log in to the Support Portal to view some of these articles.
Whe doing an AHV update I noticed that when going into or out of Maintenance mode some of the VMs show as VM and some have their actual name - why is this?
I’d like to sort my powered off VMs under AHV by last date they were powered on?Can you share your thoughts on how to display through AHV?
Hello All,I tried to install AOS 6.5.x with foundation 5.3.2 and nutanix but I faced from some issues. please find the below logs for reference2023-02-16 09:10:32,277Z INFO [1081/2430] Hypervisor installation in progress2023-02-16 09:11:02,318Z INFO [1111/2430] Hypervisor installation in progress2023-02-16 09:11:32,358Z INFO [1141/2430] Hypervisor installation in progress2023-02-16 09:12:02,371Z INFO [1171/2430] Hypervisor installation in progress2023-02-16 09:12:32,413Z INFO [1201/2430] Hypervisor installation in progress2023-02-16 09:12:32,595Z WARNING Hypervisor installation takes longer than usual2023-02-16 09:13:02,453Z INFO [1231/2430] Hypervisor installation in progress2023-02-16 09:13:32,486Z INFO [1261/2430] Hypervisor installation in progress2023-02-16 09:14:02,525Z INFO [1291/2430] Hypervisor installation in progress2023-02-16 09:14:32,563Z INFO [1321/2430] Hypervisor installation in progress
I was intially told I could use Move to take VMs from old Cluster to New Cluster but alas Move does not move from AHV to AHV, this I find very bizarre as that should be more simple than moving esxi to ahv!So, I have been told to use Data Protection failover - this seems complicated and needs VMs to be stopped and restarted - I came across Live Migration using Leap - ah ha that looks like a good option but still apears complex. Do I really need 2 Prism Central instances? Could I add the second cluster to my exisiting Prism Central then use the Migrate function to move some VMs or is the limitation of different harware still the gotcha?Does anyone have any ideas how I can do this “easily” and “quickly” my boss ain’t a happy man as information from Nutanix has been misleading or downright incorrect.Any help appreciated.
I am curious why when sizing a Splunk workload and increasing the performance profile for the indexers it doesn’t increase the ingest rate per indexer rather than keeping it at 100 GB. By not doing this, it keeps the same number of indexers Nutanix thinks it needs even though each indexer has more CPU/RAM. I could be missing something simple here, but I would like to hear other thoughts.Capacity Planning Manual: Summary of performance recommendations
Hello, after a firmware upgrade, one host is locked DOWN in maintenance mode :CVM: 192.168.131.132 DownI can run a command to exit maintenance mode but it is not working and it is "Removed from metadata store" :nutanix@:~$ ncli host edit id=7 enable-maintenance-mode=falseId : …Hypervisor Address : 192.168.131.122Host Status : NORMALOplog Disk Size : 394 GiB (423,054,278,649 bytes) (3.9%)Under Maintenance Mode : false (ncli_manual)Metadata store status : Node is removed from metadata store...So I tried to recover it but the script fails :nutanix@:~$ python /home/nutanix/cluster/bin/lcm/lcm_node_recovery.py 192.168.131.122Recovering node 192.168.131.122Checking if the node 192.168.131.122 is in phoenixCurrent node status host Node 192.168.131.122 out of phoenix modeBringing host None out of maintenance mode Successfully put host None out of maintenance mode Bringing CVM 192.168.131.122 out of maintenance modeTraceback (most recent call last): File "/home/nutanix/cluster/bin/lcm/lcm_node_
Hi there, I need some explanation about fault tolerance in a particular situation.The configuration is one 4-nodes block with each nodes configured with 2 SSDs and 4 HDDs. Also Fault Tolerance FT=1 is configured.In that case, how many disk/node are protected ?
Below are new knowledge base articles published on the week of February 5-11, 2023.KB 13748 - Accessing the IPMI on-site KB 13807 - NCC Health Check: witness_fault_domain_check KB 14190 - SQL Server Standard Edition is not able to utilize all the CPUs assigned to the VM KB 14218 - Nutanix Database Service - VM provisioning fails if AD gMSA whitelist is used KB 14229 - After upgrading to pc.2022.9, issue with upgrade of microservice infrastructure OR MSP base services KB 14233 - Enabling Distributed Autorid on systems upgraded to Nutanix Files 4.2 or later KB 14243 - Nutanix Database Service operations fail with invalid authentication credentials KB 14278 - Alert - A20032 - Containers are marked for removal KB 14284 - NDB - Clone Database Failed KB 14295 - NDB Postgres DB Server provisioning fails with "Failed to configure storage for database" with 1Gb memory profile KB 14306 - After upgrading Prism Central from pc.2022.6 to pc.2022.9, the Security Dashboard might fail to runNote: You
Hi!We are running about 100 vm’s in our cluster - so find out a specific information manually for each vm is pretty hard.I want to get the information which vm’s are stored in a specific storage container with acli on cvm.I know that i could perform acli vm.get <vm-name> and look for the source_nfs_path of the vmdisks but it is way too much effort for 100 vm’s. Another way would be establishing a connection via SSH to /storage-container/.acropolis/vmdisk of of the cvm but i only see the vmdisk-uuids there, not the vm names i need.Is there any possibility, maybe with a for loop on each vm and a grep filter?
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.