Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
The dynamic Inventory of a Hosts and the creation of Service Checks is one of the benefits of the checkmk opensource tool CheckMK. This Tool could be installed as a single Installation but it is also part OMD or Openitcockpit, which are using checkmk as an extension In my Homelab i will give a short overview of the use of checkmk with Openitcockpit with a nutanix 3 node cluser.Requirements: Enable SNMP in Prsim Centrala) Port 161/UDPb) Ein Username und Kennwörter fuer SHA und AESIf the Requirements are finished just create a new Host in Openitcockpit with Default Values. Ping is the standard Check first then. AFTER the Creation and Export of the Config we chood CHECKMK_DiscoveryWe choose SNMP V3 as Discovery Method and type in the Credentials which we prepared a step above in Prism Central.The FIRST Discovery tooks some time! If no result is shown, the credentials may we wrong or MD5 is set instead of SHA in the Auth section! If its working we get a bunch of discoverd services now:We g
Below are new knowledge base articles published on the week of September 27-October 3, 2020.KB 8607 - NCC Health Check: inconsistent_file_groups_check KB 9072 - NCC Health Check: file_server_task_stuck_check KB 9491 - NCC Health Check: container_on_removed_storage_pool KB 9728 - NCC Health Check: nearsync_stale_staging_area_check KB 9902 - NCC Health Check: vm_updates_disabled_check KB 10050 - Prism Central throwing "ErrorCode: 4" on VM update operations KB 10067 - Karbon - Private Registry Addition Troubleshooting KB 10071 - Move - Service start timed out error on Hyper-V KB 10077 - Alert ID 130096 - Failed to Recover NGT Information for VM KB 10079 - Alert ID 130095 - Failed To Recover NGT Information KB 10081 - Alert - A21014 - CassandraWaitingForDiskReplacement KB 10091 - Ubuntu 20.04 Server installation failes in guest VM running on AHVNote: You may need to log in to the Support Portal to view some of these articles.
Nutanix Clusters Console receives the results of the status checks performed by AWS to see the status of the instances.Status checks on AWS are performed every minute, returning a pass or a fail status. If all checks pass, the overall status of the instance is OK. If one or more checks fail, the overall status is impaired. There are two types of status checks, system status checks, and instance status checks. System status checks monitor the AWS systems on which instance runs. Instance status checks monitor the software and network configuration of individual instances.If Nutanix Orchestrator detects that AWS has marked system status or instance status of an instance impaired, following WARNING message will be seen in the Notification Center of the Nutanix Clusters Console: More information on - http://portal.nutanix.com/9704 What does Nutanix Cluster On AWS mean :- Nutanix Clusters provides a single platform that can span private and public clouds but operates as a single c
Nutanix Move is a very versatile tool which is used to migrate existing VMs, from various storage sources and hypervisors, into a Nutanix infrastructure. Many users have already enjoyed its functionality and ease-of-use.Often, users will migrate multitudes of VMs into their new Nutanix infrastructure using Move, and then either delete or forget about the Move VM itself thereafter. Later, when more VMs are sought to be migrated, users will either reinstall Move again or attempt to leverage the Move VM that remained from the previous migration.For those users leveraging an existing Move VM within their infrastructure and, as it is generally a good idea to use the latest version of Move when possible, it is possible to upgrade the Move VM to the latest available version right from its own dashboard (and, even by CLI if desired).You can find more information regarding upgrading Move from within the Online Upgrade section of the Move User Guide.
Hardening is the process of securing a system by reducing its surface of vulnerability, which is larger when a system performs more functions; in principle a single-function system is more secure than a multipurpose one. Reducing available ways of attack typically includes changing default passwords, the removal of unnecessary software, unnecessary usernames or logins, and the disabling or removal of unnecessary services. There are various methods of hardening Unix and Linux systems. This may involve, among other measures, applying a patch to the kernel such as Exec Shield or PaX; closing open network ports; and setting up intrusion-detection systems, firewalls and intrusion-prevention systems. There are also hardening scripts and tools like Lynis, Bastille Linux, JASS for Solaris systems and Apache/PHP Hardener that can, for example, deactivate unneeded features in configuration files or perform various other protective measures. We can implement Security Hardening features for Nutani
Application monitoring provides visibility into integrated applications by collecting application metrics using Nutanix and third-party collectors, providing a single pane of glass for both application and infrastructure data, correlating application instances with virtual infrastructure, and providing deep insights into applications performance metrics. The monitoring integrations dashboard allows you to view information about select applications, such as SQL Server instances, running in the cluster. Before you decide to enable this, there are some pre-requisites that you need to follow. To learn more about the prerequisites, click here and for more information about Application Monitoring, check the Application Monitoring Guide.
When a drive on a host (SSD or HDD) experiencing a recoverable errors, warnings or a complete hardware failure, the stargate service will mark the disk as bad. The following can be observed when a disk fail occurs. Disk is shown as a red or a plain grey in prism A critical error in Prism stating the disk went bad. Troubleshooting steps Identify the problematic disk in Prism. Check the Prism web console for the failed disk. In the Diagram view, you can see red or grey for the missing disk. Check the alerts in the Prism web console for the disk alerts, or use the following command from any of the working CVMs in the cluster to check for disks that have generated the failure messages. ncli alert ls Check to see if the disk is being recognized by the black plane. Execute the following command from the cm of the node that shows disk fail list_diks Check to see if the disk is mounted on the node. df -h Check for offline disks using NCC check disk_online_check. ncc health_checks
Using the RestAPI, we are able to pull around 300 metrics from Nutanix. Do you have a document which explains what all these metrics are, similar to the page we have for the metrics available via SNMP (https://portal.nutanix.com/page/documents/details?targetId=Web-Console-Guide-Prism-v511:man-nutanix-mib-r.html)
To start this topic, it is worth to mention the difference between the data size units. There are GBs (gigabytes) and GiBs (gibibytes). The difference is the following:GB = 1000 MBGiB = 1024 MiBSo, the difference is in binary vs decimal representation. It can be confusing, because a lot of people have never heard of gibibytes and always thought that GB is 1024 MB. In fact, it is not exactly true Why it may be confusing in Nutanix running on ESXi hypervisor?Nutanix is using MiB, GiB, TiB, etc as data size units. When you create a storage container, you get the size in binary units. For the example, i have created 2 containers with advertised capacity of 1000 GiB and 1024GiB:We can see, that the 1000 GiB container is 0.98 TiB is total size, because 1 TiB=1024 GiB.However, when we go to the vCenter, we can see a different picture:vCenter reports that the 1000 GiB container is 1000 GB in size and 1024 GiB container is 1TB. But 1 TB is 1000 GB! The problem here is that the vCenter shows teb
Below are the top knowledge base articles for the month of September 2020.KB 4116 - NX Hardware [Memory] – Alert - A1187, A1188 - ECCErrorsLast1Day, ECCErrorsLast10Days KB 7503 - NX Hardware [Memory] – G6, G7 platforms - DIMM Error handling and replacement policy KB 4141 - Alert - A1046 - PowerSupplyDown KB 1540 - What to do when /home partition or /home/nutanix directory is full KB 4158 - Alert - A1104 - PhysicalDiskBad KB 6970 - PE-PC Connection Failure alerts KB 1113 - HDD/SSD Troubleshooting KB 4519 - NCC Health Check: check_ntp KB 4409 - LCM: (LifeCycle Manager) Troubleshooting Guide KB 2090 - AHV | Host Networking KB 2473 - NCC Health Check: cvm_memory_usage_check KB 4272 - Alert - A6516 - Average CPU load on Controller VM is critically high KB 1863 - NCC Health Check: sufficient_disk_space_check KB 3741 - NGT: Nutanix Guest Tools Troubleshooting Guide KB 3784 - Alert - A1030 - StargateTemporarilyDown KB 7386 - NCC Health Check: power_supply_check KB 6945 - How Upgrades Work at N
Hey guys,I’m new at scripting/coding with python, I am trying to make a script that will automate adding a proxy, proxy whitelist and remove whitelists. This is what I have so far, but whenever it runs I get an error of “Error: Argument 'http-proxy' is specified multiple times.” import subprocessimport sysadd_whitelists = "ncli http-proxy add-to-whitelist target-type=IPV4_ADDRESS target="whitelist_Ip = ["192.168.1.1", "200.200.1.1"]login = "nutanix@192.168.1.50"add_pxy = "ncli http-proxy add name=NPX5-proxy address=proxy.npx5.local port=8080 proxy-types=HTTPS"list_proxy = "ncli http-proxy get-whitelist"#ssh connectionssh = subprocess.Popen(["ssh", "-i .ssh/id_rsa", login], stdin =subprocess.PIPE, stdout=subprocess.PIPE, stderr=subprocess.PIPE, universal_newlines=True, bufsize=0)# Send ssh commands to stdinssh.stdin.write(add_pxy)for ip in whitelist_Ip: ssh.stdin.write(a
Hi, I’m looking to understand the possible values for group_member_attributes in the POST /v3/groups request. I’ve gotten some examples from debugging the Prism UI requests (e.g. “project_name”, “vm_name”), but is there a full list published somewhere? I’d looked in the API docs and had a search of the forum but couldn’t see group_member_attributes detailed - apologies if I’ve missed it. Thanks in advance.
Hi all, I tried to get the kubeconfig renewal made by Ansible with the use of the Karbon API. All is fine except for the format of the kubeconfig file. I have a lot of \n characters and some additionnal information. Is there a way of getting a correct kubeconfig file without all these or we have to clean the file in order to get it work ?For example, this is the content of the file : {"kube_config":"# -*- mode: yaml; -*-\n# vim: syntax=yaml\n#\napiVersion: v1\nkind: Config\nclusters:\n- name: xxxx-xxxx-xxxx-\n cluster:\n server: https://xx.xx.xx.xx:443\n certificate-authority-data: LS0tLS1CRUdS0tCk1JSURyVENDQXBXZ0F3SUJBZ0lVWkJuMkxpWG13aURvbWVxQU5wWmY3emxoaEpjd0RRWUpLb1pJaHZjTkFRRUwKQlFBd2JURUxNQWtHQTFVRUJoTUNQ2tOaGJHbG1iM0p1YVdFeEVUQVBCZ05WQkFjVApDRk5oYmlCS2IzTmxNUkF3RGdZRFZRUUtFd2RPZFhSaGJtbDRNUlV3RXdZRFZRUUxFd3hKYm1aeVlTQkxZWEppCmIyNHhEVEFMQmdOVkJBTVRCSEp2YjNRd0hoY05NakF3TmpJeU1Ea3dNVEF3V2hjTk16QXdOakl3TURrd01UQXcKV2pCdE1Rc3dDUVlEVlFRR0V3SlZVekVUTUJFR0ExVUVDQk1LUTJGc2FXWnZjbTV
A traditional Nutanix cluster requires a minimum of three nodes, but Nutanix also offers the option of a one-node or two-node cluster for ROBO implementation. Both of these have a minimum requirement for the CVM to be allocated with 6 vCPUs and 20 GB memory.Specifically for two-node, there have been several improvements in 5.10 to ensure optimal health of the cluster.Customers should upgrade to a minimum version of AOS 5.10.7 to avoid encountering any of the potential service impacting issues present in earlier releases of AOS.The table provides a summary of improvements in the document here-https://portal.nutanix.com/page/documents/kbs/details?targetId=kA00e0000009DaXCAUMore information about two node cluster can be found out -https://portal.nutanix.com/page/documents/details?targetId=Web-Console-Guide-Prism-v5_18:wc-cluster-two-node-c.htmlROBO implementation meaning - Short for remote office, branch office ROBO is a term used to refer to any off-site office that connects to the organ
Era is a database as a service (DBaaS) that automates and simplifies database administration, brings one-click simplicity and invisible operations to database provisioning and life-cycle management. Era enables database administrators to provision, clone, and refresh the database clones to any point in time. Era allows administrators to define standards for their database provisioning needs with end-state driven functionality that includes High Availability (HA) database deployments. Era automates and simplifies the operations such as provisioning of databases and copy data management.Era enables you to easily provision database environments (either production or otherwise) on your Nutanix clusters. Also, you can only provision the database server VM that hosts a database, so that you can later create or clone databases on that database server. Era provisioning service includes the following components: Database engines: Custom software images that are tailor-made to enterprise needs.
Objective:Create a virtual/physical separation between different kinds of CVM traffic. Each type of traffic can be on different vLANs (virtual) or a completely different physical switch. Different type of CVM traffic are:Management: Prism, SSH, Rsyslog, SNMP, PE-PC etc. Any communication that requires the default gateway. Backplane: Mainly CVM CVM communication that happens between cluster services. Host Host and Host CVM traffic also comes under this category. Service: This is a user defined traffic type where-in he can choose a particular AOS feature to be separated out of other traffic types. RDMA: This is a type of Service traffic but only limited to Stargate service.Solution Summary:The way to achieve these separations is by creating a new interface (vNIC) on CVMs for each traffic type. By default all the traffic types happen over management interface (eth0) and user can decide to segment other traffic types to new CVM interfaces one at a time. RDMA is a special segmentation ty
Hello,have you guys ever configure a 4 SFP + Nutanix cluster with AHV? i understand that the 4 SFP + will be configured as active-passive in a same bond... but only 1 SFP should be used. There is a better way to configure network on AHV to improve and use the other SFP+??? i mean, for example, separate the traffic between user VMs and CVM (storage traffic)
Can VMs keep uuid when migrated through Xi Leap and back to original cluster?
Below are new knowledge base articles published on the week of September 20-26, 2020.KB 8911 - Alert - Flow Visualization Statistics Collector Service Restart Detected - ConntrackStatsCollectorServiceRestart KB 9622 - NCC Health Check: pc_fs_inconsistency_check KB 10023 - Push images from one PC cluster to another PC cluster KB 10036 - Windows Portable Foundation Application may fail to start without any error on Windows 10 1809 and above. KB 10048 - Metric-server cannot scrape metrics on K8s cluster KB 10049 - Prism Central throwing "BEARER_TOKEN_BAD_SIGNATURE" error KB 10056 - Troubleshooting common issues when discovering a node’s BMC information for out-of-band power management in X-RayNote: You may need to log in to the Support Portal to view some of these articles.
I tried to create a boot UEFI with URL https://github.com/abbbi/nutanix_uefiWhat information should I put for the image ce-2019.11.22-stable? menuentry 'Nutanix Community Edition AHV (4.4.77-1.el7.nutanix.20190211.279.x86_64) 7 (Core)' --class nutanix --class gnu-linux --class gnu --class os --unrestricted $menuentry_id_option 'gnulinux-4.4.77-1.el7.nutanix.20190211.279.x86_64-advanced-4d8b0f1e-e014-4058-a050-c6d2ed188094' { load_video set gfxpayload=keep insmod gzio insmod part_msdos insmod ext2 if [ x$feature_platform_search_hint = xy ]; then search --no-floppy --fs-uuid --set=root 4d8b0f1e-e014-4058-a050-c6d2ed188094 else search --no-floppy --fs-uuid --set=root 4d8b0f1e-e014-4058-a050-c6d2ed188094 fi linuxefi /boot/vmlinuz-4.4.77-1.el7.nutanix.20190211.279.x86_64 root=UUID=4d8b0f1e-e014-4058-a050-c6d2ed188094 ro crashkernel=128M rhgb quiet hugepages=0 intel_iommu=on,igfx_off iommu=pt elevator=noop vga=791 vfio_iommu_type1.allow_unsafe_interrupts
Nutanix infrastructure customers often suffer performance issues on their database, and call in to Nutanix Technical Support to help them resolve them. One aspect I have observed with database customers, and which impacts their performance, is the placement of the VMDKs on SCSI Controllers. Assuming that there are four VMDKs serving the database VM, they might have configured the four VMDKs as follows. This is an extract from the ESXi Host which the VM is running on:[root@esxasrk63u35:/vmfs/volumes/9f254e88-e3c751d/pqrsd00][root@esxasrk63u35:/vmfs/volumes/9f254e88-e3c751d/pqrsd00] cat pqrsd00.vmx | grep vmdk -A3scsi0:0.fileName = "pqrsd00.vmdk"sched.scsi0:0.shares = "normal"sched.scsi0:0.throughputCap = "off"scsi0:0.present = "TRUE"--scsi0:1.fileName = "pqrsd00_1.vmdk"sched.scsi0:1.shares = "normal"sched.scsi0:1.throughputCap = "off"scsi0:1.present = "TRUE"--scsi0:2.fileName = "pqrsd00_2.vmdk"sched.scsi0:2.shares = "normal"sched.scsi0:2.throughputCap = "off"scsi0:2.present = "TRUE"-
So you have decided to relocate your Nutanix cluster to a different data center. Here are a few things to consider and a brief overview of steps to follow for seamless transition. Caution: This information is only provided to serve as a guide to plan your move. Please engage Nutanix Support if you have any concerns or questions following this process. Before you decide to move: Consider the possibility of incorporating the existing IP address schema into the new infrastructure by reconfiguring the router and switches instead of Nutanix nodes and CVMs. If that is not possible, proceed with this guide. Before you unplug everything:Refer to these guides for the procedure.Doc 1 (CVMs) CHANGING THE CONTROLLER VM IP ADDRESSESDoc 2 (AHV hosts) CHANGING THE IP ADDRESS OF AN ACROPOLIS HOSTDoc 3 (IPMI) CHANGING AN IPMI IP ADDRESS A few things to note:1. Since the cluster is being relocated and the new network will not be able to communicate with the old network, you will need to run through so
If there was a gflag applied in the past on cluster and now you dont have a track on the details, the same can be found out by running this NCC check ncc health_checks system_checks gflags_diff_checkThe output will show you the list of non-default gflag currently on the system, more details here -NCC Health Check: gflags_diff_check
Very frequently we receive an alert just before an upgrade that cluster doesn't have enough space to download the binary files (). If that is happening to you, there are ways to clean the partition below the threshold (75%). Some of them include below Cleaning Old ISOs and Software Binaries ( For old AOS, NCC, Foundation versions) Checking Removing Old Logs (Log files shared for support cases if not cleared later, occupy space) Cleaned up files from the approved directories but still see high usage in /home?At this time, open a support case to identify any other underlying issues or deep cleaning on the nodes with the help of Nutanix Engineer. More details on Nutanix KB below https://portal.nutanix.com/page/documents/kbs/details?targetId=kA0600000008dpDCAQ
Nutanix has worked closely to build integrations between ServiceNOW’s SaaS offering andNutanix infrastructure with the following areas of focus:ServiceNOW CMDB integration for Nutanix hyper-converged infrastructure discovery and modeling Nutanix Alert Integration with ServiceNOW Events and Incidents and remediation via XPlay ServiceNOW Request Management and Fulfillment for Services Hosted on Nutanix via CALM plug-in For More Details: https://www.nutanix.com/blog/nutanix-integration-with-servicenow
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.