Visualizzazione post con etichetta SSP. Mostra tutti i post
Visualizzazione post con etichetta SSP. Mostra tutti i post

venerdì 20 marzo 2026

[SSPi] - Backup Process fail

Issue


The SSPI backup process fails and displays the following error message: “Management cluster backup failed. Please check logs for more details.” as shown below ..
management cluster backup failed, please check logs for more details.
Command '/usr/local/bin/clusterctl move --to-directory /image/br/sspi/mgmt-cluster/move/ --config /config/clusterctl/1/clusterctl-init.yaml --namespace ssp-rm-w01 --v 5' execution is failed. Output: Using configuration file="/config/clusterctl/1/clusterctl-init.yaml" 
.....
.....
ExecuteC github.com/spf13/cobra@v1.9.1/command.go:1148 github.com/spf13/cobra.(*Command).
Execute github.com/spf13/cobra@v1.9.1/command.go:1071 sigs.k8s.io/cluster-api/cmd/clusterctl/cmd.Execute sigs.k8s.io/cluster-api/cmd/clusterctl/cmd/root.go:113 main.main sigs.k8s.io/cluster-api/cmd/clusterctl/main.go:27 runtime.main runtime/proc.go:272 runtime.goexit runtime/asm_amd64.s:1700 , 
Error:exit status 1 


Solution


Disclaimer: Some of the procedures described below may not be officially supported by VMware. Use it at your own risk. I suggest to contact broadcom support to troubleshoot this issue.

Googling around, I found the KB that indicates exactly my error "Management cluster backup failed. Please check logs for more details.", and which tells me to contact support.

I tried to investigate by connecting via SSH to the SSPI and typing ...
df -h
The root "/" directory is 100% used, which is not good at all, regardless.
In my cases, I emptied the root folder by compressing the "log" folder (13 GB) (which was under the "/root" directory) previously used for troubleshooting.
Then I tried to manually launch the backup and it worked.


That's it.

domenica 1 marzo 2026

[SSP 5.1.0] vDefend SSP Backup Failures: Stuck at 60%

Issue


If you manage a VMware vDefend Security Services Platform (SSP) environment, you know how critical automated backups are. Recently, I ran into a frustrating issue where scheduled backups started failing mysteriously. What initially seemed like a simple user and misconfigured settings issue, “network timeout” turned out to be a bizarre problem involving lost encryption keys and crashing containers.
Looking on the SSP UI at the System > Backup and Restore dashboard, I noticed a long list of failed backup jobs. Hovering over the failure status yielded a very generic message: "operation timed out, please contact administrator".
I tried manually forcing the backup, but it seems to stuck at 60% and then time out after a while.
The generic timeout error on the Backup and Restore dashboard


Solution


My first instinct was to check the SFTP credentials, re-type the password, SAVE and seemed fine.
Googling around, I found this interesting link "Backup and Restore goes into Failed State" which helped me in the final resolution.

Since the UI wasn't giving me enough technical details, it was time to drop into the CLI. I logged into the SSPI appliance via SSH as sysadmin to check what was happening at the Kubernetes pod level.

I ran a quick ...
kubectl get pods -n nsxi-platform | grep backup
... and immediately spotted the problem. The backup job pods weren't just timing out; they were actively crashing, throwing Error and OOMKilled statuses.
To understand why the pods were in this state ("OOMKilled", "Error"), I checked the logs of both pods....
k logs job-backup-4fv4k-5gq5s -n nsxi-platform
Scrolling through the container logs, I could see ... the system passed a malformed argument to the OpenSSL encryption command.
The log explicitly stated:
Error executing command: [openssl enc -aes-256-cbc -salt -in ... -out ... -k ******]. 
Error: fork/exec /usr/bin/bash: argument list too long.
Scrolling through the container logs, I could see the backup job actually doing its work: generating the .obj and .dump files from the Postgres database and config files.
However, the process abruptly halted right after the ...
<INFO> ... Started encrypting backup bundle
... message.


Conclusion

How did I fix it? I re-entered the password and retyped the passphrase to ensure the openssl process received a valid string, allowing the container to properly encrypt the dump files and upload them to the SFTP server.

The mystery was solved. The "timeout" in the UI was just a side effect. The real issue was that the missing encryption key caused the openssl command to fail spectacularly (resulting in an argument list too long OS error), which crashed the container pod.

It appeared the platform had somehow lost track of the encryption key. The configuration was stuck in a weird state where trying to "Reuse prior passphrase" wasn't working.

Now starting the backup in manual mode, it completes successfully.


That's it.