Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
Hello, I am new and I am interested in whether Nutanix acropolis supports Atlassian products Bamboo, Bitbucket, Jira? Does anyone have experience with their containers on the Nutanix platform?
I am testing Postman for API calls and one of the errors I am getting is 502 - Oops Server Error, which is related to the Apache portion. Question is, which log contains this? I’ve checked all the logs in ~/data/logs/ and can’t find that response.
Here’s my problem. Node A & B are synchronizing to the wrong host. Node C is synchronizing to the right host. This is the behavior I see on the CVMs. When I do “hostssh ntpq -pn”, the hypervisors (AHV) are reporting the correct ntp server. How do I bring nodes A & B back inline to what they should be? I tried manually correcting their ntp.conf files with the correct IP and restarted the ntpd service. No change and eventually the ntp.conf files changed to the wrong settings. Not sure how to wrestle this one to the ground.
Below are new knowledge base articles published on the week of June 14-20, 2020. KB 8507 - Cosmetic high latency spikes may be observed in Prism at time of low IOPS on a cluster KB 8932 - NCC Health Check: pc_vm_resource_resize_check KB 9423 - Beam Cost Governance for Nutanix On-Prem stops reporting cost analytics post Cluster upgrade to AOS 5.15 KB 9488 - AHV | CentOS/RHEL 6.8 may hang durin the boot when tboot package is used KB 9496 - Prism Central UI showing license expired red banner even after disabling prism pro features KB 9503 - Genesis crashing on a ESXI Node | No services are up | pyVmomi issue | ParserError: 'xml document KB 9516 - Using Nutanix objects as static web server for LCM dark sites KB 9519 - Nutanix Move | How to Configure HTTP(s) Proxy on Move KB 9534 - Hyper-V: Hyper-V 2019 fails during Foundation with "InstallerVM timeout occurred, current retry 0" KB 9538 - Prism logins fail if service account AD permissions are default KB 9545 - Copier reports inc
Hello guys, currently i deployed a new hpe hardware Hypervisor version is ESXi VERSION 6.7.0 I need a vcenter to deploy some VMs in Nutanix, so i downloaded the vcenter applicane and tried to run the setup with admin rights. I have currently no dns server, so i used the gateway. NTP server is currently not available, because the environment is completly new, so i used the Time-synchronisation: Synchronize time with the ESXi hostSynchronize time with NTP servers Its activated with http://de.pool.ntp.org/; but is only temp. The vcenter appliance setup 1 is running fine, the step 2 is very slow and need much time. After 2hours on 50% i canceled it. The error: No file found matching /etc/vmware-vpx/vc-extn-cisreg.prop cmd "/usr/bin/python /usr/lib/vmidentity/tools/scripts/lstool.py list --url http://localhost:7080/lookupservice/sdk" timed out after 60 seconds due to lack of progress in last 30 seconds (0 bytes read) but why? Should I try another vcen
Hi, A customer moved his Nutanix Cluster (with Hyper-V) from a DC to another, after powering the Nodes up, the IPs of all Hyper-V host and CMs released, I logged it locally to Hyper-V hosts and configure the internal IP (192.168.5.1/28) and the external IP same like before shutting the cluster down. I repeated the previous step with CVMs, I went through cd/etc/sysconfigs/ and edit network-scripts file and added the the external IP in the eth0 and the internal one (192.168.5.2/28) in eth1. Now Hyper-V FC is working fine but cannot start VMs due to the Nutanix cluster issue, whenever I tried to start cluster from any CVM, I get this message “WARNING genesis_utils.py:1211 Failed to reach a node where Genesis is up. Retrying” Is there any way to fix this issue or to repair cluster configuration without disrupting existing data? Thanks in advance
esxi 6.5 u2 vSwitch configured as 1 active uplink and 1 standby uplink on esxi nutanix. I didnt understand why the standby adapter in which scenarios. I wonder why and in what scenarios it is used. can i use both 10 GB network adaptor active-active mode in Virtual Standard Switches (VSS).
Ques- There is 5 node cluster with 200 TB raw disk data, administrator want to enable eraser coding to obtain more space. What will we the usable space after enable eraser coding Options – 100, 125, 150, 175 TB ? Please share the calculation formula ? Ques2 - Administrator want to perform DR test after 6 months, which snapshot he should use? Option – latest snapshot, oldest one
In this topic I will share how to configure an SMTP server on your Nutanix cluster. This can be done from the Prism UI. This is pretty simple and straight forward. Below is the link shared, that discusses how to configure a SMTP server- https://portal.nutanix.com/#/page/kbs/details?targetId=kA032000000TTWtCAO After the SMTP server is created, the email alerts should be tested if they are being sent from the cluster. SMTP always uses port 80 and 8443, so make sure that these firewall ports are open. You could send test emails to verify correct SMTP configuration. This can be done using the KB2773. I will attach a link to the KB below. https://portal.nutanix.com/#/page/kbs/details?targetId=kA03200000097qFCAQ Hope this piece of information is useful!
Is it possible for a single bridge can have multiple bonds? in which usecase it will be applicable? -Jai
Below are new knowledge base articles published on the week of June 7-13, 2020. KB 9456 - Alert - A400114 - PolicyEngineServiceDown KB 9468 - Different Docker hosts can see volumes created with Nutanix DVP KB 9478 - How to clear stuck LCM inventory tasks ? KB 9480 - Windows 10, version 2004 or Windows Server, version 2004 VMs may fail to boot on AHV KB 9487 - AHV | VM update operations initiated from Prism Central may fail with "Entity CAS version mismatch" error due to missing machine_type attribute KB 9492 - Move VM IP address may change post deployment KB 9494 - Disks from ISCSI volumes change drive letters on Windows VMs after upgrade AHV 2016* to 2017* Note: You may need to log in to the Support Portal to view some of these articles.
Hi all, A newbie question. It seems I still have an old crash dump directory on one of my AHV hosts. A ls -lahtr /var/crash shows the single directory from back in April. At the time the issue was resolved, and the faulty DIMM that was causing the issue was replaced. That said, clearly the dump file was not removed. In terms of cleaning this up, is it OK to delete the dump directory within the /var/crash directory and then rerun the ncc health check. Or is there a better method for clearing crash dumps from Nutanix clusters. Many thanks, Rob
Please answer the below questions : When one cvm goes down(in a 10 node cluster with RF3) for 20 mints, then Guest VM’s new write and read IO would be served by anther's CVM and all these will be traveling across 10 g network. A : Will those new IOs served from local copy via another CVM or IOs will be served from the replica copy B: Will cluster starts to build a new replica to accommodate the missing copy or RF3 Ques 2 - When new Write IO request comes, first it will write the data on Oplog then synchronously send to other CVM’s Oplog. All host in clusters having 2 SSD and 6 HDDs. Will Write IO process by both SSD in every host or only one SSD is hsving Oplog partition active at the same time ? I think there is only one oplog per CVM/host whether the host is aving all flash drive or 2 SSD and remaining HDD. But I am not sure about this statement, Please clarify..
Is it possible to recover data from an overwritten entity - does the overwritten VM or its storage still exist anywhere? I'm asking as I recovered a VM from a backup but realised in hindsight that the logs from the VM in its failed state would be useful for diagnosing the issue. Thanks in advance.
What is the DIMM error? A memory error is an event that leads to the logical state of one or multiple bits being read differently from how they were last written. For example, If 1 was written in a memory cell and while reading the same memory cell, it returns 0. Memory errors can be classified into two types: Soft errors, which randomly corrupt bits but do not leave physical damage. Soft errors are transient in nature and are not repeatable. Soft errors can be because of electrical or magnetic interference (e.g. due to cosmic rays, alpha particles, leakage, random noise). Hard errors, which corrupt bits in a repeatable manner because of a physical/hardware defect or an environmental problem. Hard error can also occur if DIMM is not seated properly. All memory systems in use in servers today are protected by error detection and correction codes. These server machines employ error correcting codes (ECC), which allows the detection and correction of one or m
First, let us understand what NTP (Network time protocol) is. An NTP server is a time server that is used to keep/sync the time in your cluster. An NTP server can be public or private depending on the strictness of your environment. To know how to configure NTP in your Nutanix cluster, take a look at- https://support-portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v5_16:wc-system-ntp-servers-wc-t.html After the NTP server is configured, the genesis leader becomes the NTP leader, which means that the genesis leader is syncing time to the NTP server and other CVMs are syncing time with the genesis leader. How NTP works in AHV:- It’s as simple as it gets. The AHV hypervisor takes the same server configured on the cluster and syncs the time with it individually. There are no extra steps required to configure the NTP server on the AHV hosts. How NTP works in ESXi:- The ESXi cluster does not take the server configured on the Nutanix cluster and it needs to be
Hello, I’m working with Rest API, to “automate” some operations with Ansible, trought URI Module and Jinja2. I need to create a Project, and set permission to a specific user. How can i do this with API V3? I can add user to a Project, but i can’t set the Role.. When i check trought the Web Interface, i see the user without role. Thanks!
Below are new knowledge base articles published on the week of May 31-June 6, 2020. KB 9442 - LCM BIOS/BMC Upgrade fails when node does not respond to IPMI power reset KB 9460 - Move: ESXi-AHV migration connection limits KB 9463 - Pre-upgrade check: Hypervisor Upgrade (test_host_upgrade_versions_compatible) KB 9467 - How to enable Karbonctl in Karbon darksite environment Note: You may need to log in to the Support Portal to view some of these articles.
Hi, everyone I’m trying to use LCM to update firmware of host machines on my cluster. It all went well until the post action phase. Error message said: ‘Operation failed. Reason: LCM failed performing action reboot_from_phoenix in phase PostActions on ip address xx’. I searched the KB and found KB9177,but it’s about ‘Mixed Hypervisor cluster‘, my cluster uses solely AHV so it doesn’t applied. Anyway I still tried to follow the KB9177’s suggestion and upgraded my cluster’s foundation to 4.5.3 and retried the LCM firmware update process on another host but still got the same ‘ LCM failed performing action reboot_from_phoenix’ error. I used the workaround provided in that KB to make the affected two hosts’ CVM out of maintenance mode. It works and the cluster is back to normal. Then I logged on to the affected hosts’ IMM and found out that actually the primary IMM2 firmware has already been updated by the LCM( the backup IMM2 firmware is not upgraded), and when I go to the LCM section
Hello, for LCM updates I see that pre-checks require DRS in fully automated mode. It’s possible to perform LCM updates without DRS in fully automated mode (in the event that a customer does not have DRS in his licenses) ? Thanks Manuel
Hello. I faced with limitation of number VMs in API reply. There is option “length” which can be added to request, but max value is 500 (I got it into experiments). We have around 1000 VMs in our environment. Is there solution?
Hello, We have a cluster with 4 nodes and I will perform the update from 5.9.4.2 to 5.10.10.1. a some minutes ago, I see that my cluster have a disk space problem and the resilience board is red with the “not resilience” information. In this case I will have a problem when perform the AOS upgrade? I not have time to resolve the space issue now… Please, send me any information.
is there any way to created additional users in nutanix move? -jai
I’m trying to setup a new bridge (br1). After ssh to the Nutanix AHV, I continue with input “ssh root@xxx.xxx.xxx.xxx” to get to the CVM, it ask me for a password. Is there a way to change/ continue with a different logon name as well at that? The password I input is being denied due a different username for the CVM.
Hello, comunity I’m been trying to figure out why the configuration with RF2 & N+1 reserves more capacity than just RF2. I understand that with RF2 all the data is duplicated along the cluster and the failure of one node is tolerated. In my scenario I have a 3-node cluster with 34TiB of effective capacity, so with RF2, I would have only 17TiB, everything is clear until that point. But the extent store assuming RF2 & N+1 gives me 11.3TiB, which is 1/3 of my available capacity. So...the question is: If I reserve capacity, by tripling all the data, that is for tolerate 1 additional node failure besides the 1st one (tolerated by RF2)? A silly conclussion would be that with RF2 & N+1 the cluster is able to tolerate 2 nodes and continue operating just with one node, but I know that is not posible. So, why assuming RF2 & N+1 reserves more TiB’s than just RF2? Please, I would appreciate the help. Thanks in advanced!
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.