The purpose of this utility is to test HTTP/HTTPS reachability, alert if its down and run a config change. On recovery run another list of config changes.
Jeremy Georges - Arista Networks - jgeorges@arista.com
TCPCheck Utility
The purpose of this utility is to test HTTP/HTTPS reachability, alert if its down and run a config change. On recovery run another list of config changes.
Add the following configuration snippets to change the default behavior. Current version supports only one HOST and one port.
daemon TCPCheck
exec /mnt/flash/TCPCheck.py
option CHECKINTERVAL value 5
option CONF_FAIL value /mnt/flash/failed.conf
option CONF_RECOVER value /mnt/flash/recover.conf
option FAILCOUNT value 2
option HTTPTIMEOUT value 10
option IPv4 value 10.1.1.1
option PROTOCOL value https
option TCPPORT value 443
option USERNAME value admin
option URLPATH value /index.html
option PASSWORD value 4me2know
option REGEX value HelloWorld
option VRF value mgmt
no shutdown
Config Option explanation:
- CHECKINTERVAL is the time in seconds to check the HTTP/S Neighbor(s). Default is 5 seconds.
- FAILCOUNT is the number of times/iterations that the neighbor must fail before
declaring the neighbor is down and executing config changes. This parameter is optional. The default is 2.
- IPv4 is the address to check. Mandatory parameter.
- HTTPTIMEOUT is the time in seconds we wait for an HTTP response. Default is 20 seconds.
- PROTOCOL is either http,https. This is a mandatory parameter.
- TCPPORT is the TCP port to use. Only applicable for HTTP/HTTPS. Mandatory parameter.
- USERNAME is the username needed for HTTP request. If not set,
then no username and password is sent.
- PASSWORD is the password used for the HTTP request. If not set, then no password
is used.
- CONF_FAIL is the config file to apply the snippets of config changes. Mandatory parameter.
- CONF_RECOVER is the config file to apply the snippets of config changes
after recovery of Neighbor. Mandatory parameter.
- REGEX is a regular expression to use to check the output of the http response. Mandatory parameter.
- URLPATH is the specific path when forming the full URL. This is optional. Default is just the root '/'.
- VRF is if you want the HTTP requests to use a specific VRF. This is optional. The default VRF will be used if not set.
The CONF_FAIL and CONF_RECOVER files are just a list of commands to run at failure or at recovery. These commands should be full commands just as if you were configuration the switch from the CLI (i.e. not abbreviated commands).
For example the above referenced /mnt/flash/failed.conf file could include the following commands, which would shutdown the BGP neighbor on failure:
router bgp 65001.65500
neighbor 10.1.1.1 shutdown
The recover.conf file would do the opposite and remove the shutdown statement:
router bgp 65001.65500
no neighbor 10.1.1.1 shutdown
This is of course just an example, and your use case would determine what config changes you'd make.
Please note, this uses the EOS SDK eAPI interaction module. You do not need to specify 'enable' and 'configure' in your configuration files, because it automatically goes into configuration mode.
Blank lines and comment lines beginning with '!' in the CONF_FAIL and CONF_RECOVER files are ignored, so you can annotate them freely.
This requires EOS SDK. All new EOS releases include the SDK.
The check succeeds only when the REGEX matches somewhere in the HTTP response. Getting a TCP connection and an HTTP 200 back is not sufficient. This is the single most common source of false "down" reports, so it is worth a moment's thought.
Two things to be aware of:
The response is not followed through redirects. If the server returns a 301 or 302, the body you get
is a short redirect stub, not the page you were aiming at. A stub will essentially never contain your
REGEX, so the check will report down even though the server is perfectly healthy. Verify with curl -i
and, if you see a 3xx, point URLPATH at the target of the Location: header instead.
Modern web UIs render their content in JavaScript. If the page is a single-page application, the HTML
delivered over the wire is a near-empty shell - a <div id="root"> and a <script> tag - and the text you
see in a browser is never present in the HTTP response at all. No REGEX will match it, because the string
is produced by the browser, not the server.
The Arista eAPI explorer is a good illustration of both. On older EOS releases it lived at /explorer.html
and was server rendered. On current releases /explorer.html returns a 301 to /eapi/, and /eapi/ returns
a JavaScript shell whose only stable matchable text is <title>Arista</title>. A check written years ago as
URLPATH /explorer.html with REGEX eAPI will fail on a current switch for both reasons at once.
Always confirm what the server actually sends before setting REGEX:
bash curl -i http://10.1.1.1/eapi/
and pick a string you can see in that output.
Bear in mind what a page-content check does and does not prove. Matching a title on a static page tells you the web server is answering; it does not tell you the application behind it is functional.
Agent: TCPCheck (running with PID 14743)
Uptime: 0:11:18 (Start time: Sat May 30 17:19:58 2020)
Configuration:
Option Value
------------------- -----------------------
CHECKINTERVAL 5
CONF_FAIL /mnt/flash/failed.conf
CONF_RECOVER /mnt/flash/recover.conf
FAILCOUNT 2
HTTPTIMEOUT 10
IPv4 192.168.100.103
PASSWORD 4me2know
PROTOCOL https
REGEX Arista
TCPPORT 443
URLPATH /eapi/
USERNAME admin
VRF mgmt
Status:
Data Value
------------------------ -----------------------
CHECKINTERVAL: 5
CONF_FAIL: /mnt/flash/failed.conf
CONF_RECOVER: /mnt/flash/recover.conf
FAILCOUNT: 2
HTTPTIMEOUT: 10
HealthStatus: UP
IPv4 Address List: 192.168.100.103
PASSWORD: 4me2know
PROTOCOL: https
REGEX: Arista
Status: Administratively Up
TCPPORT: 443
URLPATH: /eapi/
USERNAME: admin
VRF: mgmt
May 30 16:48:45 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: %AGENT-6-INITIALIZED: Agent 'TCPCheck-TCPCheck' initialized; pid=12335
May 30 16:48:45 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: TCPCheck Version 3.0.0 Initialized
.
After HTTP Host goes down...
.
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Connection Timeout
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: HTTP HOST is down. Changing configuration.
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Applied Configuration changes from /mnt/flash/failed.conf
.
After Recover of HTTP Host...
.
May 30 16:49:57 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: HTTP host back up. Changing Configuration.
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_E: Enter configuration mode from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_I: Configured from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_E: Enter configuration mode from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_I: Configured from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Applied Configuration changes from /mnt/flash/recover.conf
Which method you use depends on your EOS release. The 4.26.0 boundary matters because that is where Sysdb mount profiles stopped being required, which is what the RPM existed to install.
No RPM, extension or mount profile is needed. Copy TCPCheck.py onto the switch, make it executable, and
point the daemon at it. /mnt/flash is the easiest location because it persists across reloads.
Copy the file to the switch by whatever means you prefer, for example from the CLI:
copy scp://user@server/path/TCPCheck.py flash:
Then make it executable. The agent will not start without this, and the failure is not especially obvious in the logs:
bash sudo chmod +x /mnt/flash/TCPCheck.py
Confirm it took:
bash ls -l /mnt/flash/TCPCheck.py
You are looking for the x bits, e.g. -rwxr-xr-x.
Then configure the daemon:
configure
daemon TCPCheck
exec /mnt/flash/TCPCheck.py
option IPv4 value 10.1.1.1
option PROTOCOL value https
option TCPPORT value 443
option REGEX value Arista
option CONF_FAIL value /mnt/flash/failed.conf
option CONF_RECOVER value /mnt/flash/recover.conf
no shutdown
If you edited or transferred the file from Windows, make sure it has unix line endings before starting the daemon. See Agent crash due to Windows CR LF below.
Releases before 4.26.0 still require a Sysdb mount profile, which version 3.0.0 no longer ships. For those
releases use the legacy RPM in the legacy/ directory of this repository:
legacy/TCPCheck-2.3.2-1.noarch.rpm
Install it as an extension in the normal way. It takes care of all the file requirements, including placing
the mount profile in /usr/lib/SysdbMountProfiles and the agent in /usr/local/bin. With that RPM the
daemon exec line is /usr/local/bin/TCPCheck, not /mnt/flash/TCPCheck.py.
Be aware of what you are getting: the legacy RPM is version 2.3.2, which is Python 2 only. It will not run on an EOS image that has removed the Python 2 interpreter, and it does not contain any of the fixes listed under What changed in 3.0.0. It is provided for continuity on older releases, not as a maintained alternative.
Older releases of EOS required a SysdbMountProfile, and previous versions of this project shipped one as a
second file named TCPCheck (no extension) alongside TCPCheck.py. That requirement also forced the agent
filename and the mount profile filename to match.
From EOS 4.26.0 onward, mount profiles are no longer required. The SDK mounts what an agent needs automatically. Arista's guidance is to remove any mount profile you have, because running without one uses less memory and CPU than running with one.
As of version 3.0.0 this project no longer ships a mount profile. If you are upgrading from an earlier version, delete the stale file from the switch, since an RPM upgrade will not necessarily remove a file a previous version installed:
bash sudo rm -f /usr/lib/SysdbMountProfiles/TCPCheck
With the mount profile gone, the agent filename no longer has to match anything, so exec /mnt/flash/TCPCheck.py
is fine.
If you are targeting an EOS release older than 4.26.0 you will still need a mount profile, and should use
the legacy RPM in the legacy/ directory, described above.
Version 3.0.0 is a Python 3 port. It requires Python 3 and will not run under Python 2. Versions 2.3.2 and earlier require Python 2 and will not run on EOS releases that have removed the Python 2 interpreter.
| TCPCheck version | Python | EOS releases | How to install |
|---|---|---|---|
| 2.3.2 and earlier | Python 2 | Pre-4.26.0. Tested on 4.20.1, 4.20.4, 4.20.5, 4.24.0. Requires a Sysdb mount profile. | RPM in legacy/ |
| 3.0.0 | Python 3 | 4.26.0 and later, where mount profiles are no longer used. | Copy TCPCheck.py to /mnt/flash and chmod +x |
The shebang is #!/usr/bin/env arista-python. arista-python is the EOS supplied wrapper that resolves to the
correct interpreter for the image and sets up the eossdk import path, which avoids the agent silently changing
interpreter across upgrades the way a bare python shebang does.
To confirm which interpreter the SDK bindings are installed under on a given image:
bash ls -d /usr/lib/python*/site-packages/eossdk*
Changes required by Python 3:
socket.socket(_sock=...)has been removed. In Python 3socket.fromfd()already returns a usable socket object, so the re-wrap that followedVrfMgr.socket_at()is gone. The explicitos.close()of the descriptor returned bysocket_at()is still required, becausefromfd()duplicates it.ssl.wrap_socket()was deprecated in Python 3.7 and removed in 3.12, andssl.PROTOCOL_TLSv1went with it. HTTPS now uses anssl.SSLContext. Certificate verification remains disabled, matching the previous behavior, so self-signed certificates still work. A side effect is that TLS is no longer pinned to TLSv1, so HTTPS checks now succeed against servers that have dropped it - which current EOS has.base64.b64encode()takes bytes, and sockets send and receive bytes. The basic auth header, the request, and the response are encoded and decoded explicitly.
Defects fixed along the way, which were present in the Python 2 versions too:
- The response is now read until complete rather than with a single
recv(). A single read frequently returned only the HTTP headers, because servers commonly write headers and body as separate segments, which meant a REGEX intended to match page content could fail against a perfectly healthy server. The read usesContent-Lengthwhere present, so it does not wait for the server to close the connection. - A leading CRLF is no longer sent before the request line. RFC 9112 removed the allowance for servers to ignore an empty line before the request, and hardened HTTP stacks now reject it outright.
- Socket cleanup moved into a
finallyblock. Previously cleanup only ran on the success path, so an exception betweenconnect()and cleanup leaked a file descriptor once per CHECKINTERVAL. - The error handler guarding socket creation called
close()on the socket whose creation had just failed, raisingUnboundLocalErrorand masking the real error. - CONF_FAIL and CONF_RECOVER file reads are guarded, and blank and comment lines are filtered before being
passed to eAPI. Previously a trailing newline or a comment would cause the entire config batch to be
rejected, and an empty file raised
IndexError.
Every changed region in TCPCheck.py carries a comment explaining the reason, tagged PY3 PORT: where the
change was forced by the language or standard library, and BUGFIX: where a latent defect was corrected.
Based on licensing, this is open source and this and other open-source tools are not supported directly by Arista Support. Support is best effort as it relates to extensions such as this one.
Almost always a REGEX that cannot match. See Choosing a REGEX above. Confirm with:
bash curl -i <protocol>://<IPv4>:<TCPPORT><URLPATH>
Check three things in that output, in order: the status line is a 200 and not a 3xx; the body is the page you expect and not a redirect stub or a JavaScript shell; and your REGEX string is literally present in the body.
Note that curl follows redirects with -L while this agent does not, so a curl that succeeds without -i
can easily disagree with the agent. Always use -i when comparing.
If the agent is continuously crashing it might be caused by invalid characters in the python script. In the syslog outputs or show logging similar errors could be seen:
Aug 31 02:46:41 switch1 ProcMgr-worker: %PROCMGR-6-PROCESS_TERMINATED: 'TCPCheck' (PID=7319, status=32512) has terminated.
Aug 31 02:46:41 switch1 ProcMgr-worker: %PROCMGR-3-PROCESS_DELAYRESTART: 'TCPCheck' (PID=7319) restarted too often! Delaying restart for 120.0
To troubleshoot further the agent logs should be checked with bash cat /var/log/agents/<agentName>-<pid>.log (substitute <agentName>-<pid> with the actual name of the file
If the output is something similar as below, then it means the python script has an invalid \r which is the Windows carriage return (CR LF)
cat TCPCheck-Rack1-17880
==== Output from /mnt/flash/TCPCheck [] (PID=17880) started Sep 1 15:00:00.00000 ===
/usr/bin/env: 'python\r': No such file or directory
The solution is to convert the file to unix format, this can be done locally on EOS by editing the file with vi and typing :set ff=unix, so the steps would be:
- drop down to global configuration mode and shutdown the daemon
daemon TCPCheck shutdown - go to bash by typing
bash vi /mnt/flash/TCPCheck- type
:set ff=unix - press Enter
- press Esc
- type
:wq! - type
exitto go back to EOS CLI and bring up the daemon again no shutdown
Tip: When using Notepad++ to edit files always convert them to unix format by clicking on Edit - EOL Conversion and select Unix(LF) and save the file.
BSD-3, See LICENSE file