Skip to content

Latest commit

 

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

TCPCheck

The purpose of this utility is to test HTTP/HTTPS reachability, alert if its down and run a config change. On recovery run another list of config changes.

Author

Jeremy Georges - Arista Networks - jgeorges@arista.com

Description

TCPCheck Utility

The purpose of this utility is to test HTTP/HTTPS reachability, alert if its down and run a config change. On recovery run another list of config changes.

Add the following configuration snippets to change the default behavior. Current version supports only one HOST and one port.

daemon TCPCheck
   exec /mnt/flash/TCPCheck.py
   option CHECKINTERVAL value 5
   option CONF_FAIL value /mnt/flash/failed.conf
   option CONF_RECOVER value /mnt/flash/recover.conf
   option FAILCOUNT value 2
   option HTTPTIMEOUT value 10
   option IPv4 value 10.1.1.1
   option PROTOCOL value https
   option TCPPORT value 443
   option USERNAME value admin
   option URLPATH value /index.html
   option PASSWORD value 4me2know
   option REGEX value HelloWorld 
   option VRF value mgmt
   no shutdown
Config Option explanation:
    - CHECKINTERVAL is the time in seconds to check the HTTP/S Neighbor(s). Default is 5 seconds.
    - FAILCOUNT is the number of times/iterations that the neighbor must fail before
    declaring the neighbor is down and executing config changes. This parameter is optional. The default is 2.
    - IPv4 is the address to check. Mandatory parameter.
    - HTTPTIMEOUT is the time in seconds we wait for an HTTP response. Default is 20 seconds.
    - PROTOCOL is either http,https. This is a mandatory parameter.
    - TCPPORT is the TCP port to use. Only applicable for HTTP/HTTPS. Mandatory parameter.
    - USERNAME is the username needed for HTTP request. If not set,
    then no username and password is sent.
    - PASSWORD is the password used for the HTTP request. If not set, then no password
    is used.
    - CONF_FAIL is the config file to apply the snippets of config changes. Mandatory parameter.
    - CONF_RECOVER is the config file to apply the snippets of config changes
    after recovery of Neighbor. Mandatory parameter.
    - REGEX is a regular expression to use to check the output of the http response. Mandatory parameter.
    - URLPATH is the specific path when forming the full URL. This is optional. Default is just the root '/'.
    - VRF is if you want the HTTP requests to use a specific VRF. This is optional. The default VRF will be used if not set.

The CONF_FAIL and CONF_RECOVER files are just a list of commands to run at failure or at recovery. These commands should be full commands just as if you were configuration the switch from the CLI (i.e. not abbreviated commands).

For example the above referenced /mnt/flash/failed.conf file could include the following commands, which would shutdown the BGP neighbor on failure:

router bgp 65001.65500
neighbor 10.1.1.1 shutdown

The recover.conf file would do the opposite and remove the shutdown statement:

router bgp 65001.65500
no neighbor 10.1.1.1 shutdown

This is of course just an example, and your use case would determine what config changes you'd make.

Please note, this uses the EOS SDK eAPI interaction module. You do not need to specify 'enable' and 'configure' in your configuration files, because it automatically goes into configuration mode.

Blank lines and comment lines beginning with '!' in the CONF_FAIL and CONF_RECOVER files are ignored, so you can annotate them freely.

This requires EOS SDK. All new EOS releases include the SDK.

Choosing a REGEX

The check succeeds only when the REGEX matches somewhere in the HTTP response. Getting a TCP connection and an HTTP 200 back is not sufficient. This is the single most common source of false "down" reports, so it is worth a moment's thought.

Two things to be aware of:

The response is not followed through redirects. If the server returns a 301 or 302, the body you get is a short redirect stub, not the page you were aiming at. A stub will essentially never contain your REGEX, so the check will report down even though the server is perfectly healthy. Verify with curl -i and, if you see a 3xx, point URLPATH at the target of the Location: header instead.

Modern web UIs render their content in JavaScript. If the page is a single-page application, the HTML delivered over the wire is a near-empty shell - a <div id="root"> and a <script> tag - and the text you see in a browser is never present in the HTTP response at all. No REGEX will match it, because the string is produced by the browser, not the server.

The Arista eAPI explorer is a good illustration of both. On older EOS releases it lived at /explorer.html and was server rendered. On current releases /explorer.html returns a 301 to /eapi/, and /eapi/ returns a JavaScript shell whose only stable matchable text is <title>Arista</title>. A check written years ago as URLPATH /explorer.html with REGEX eAPI will fail on a current switch for both reasons at once.

Always confirm what the server actually sends before setting REGEX:

bash curl -i http://10.1.1.1/eapi/

and pick a string you can see in that output.

Bear in mind what a page-content check does and does not prove. Matching a title on a static page tells you the web server is answering; it does not tell you the application behind it is functional.

Example

Output of 'show daemon' command

Agent: TCPCheck (running with PID 14743)
Uptime: 0:11:18 (Start time: Sat May 30 17:19:58 2020)
Configuration:
Option              Value
------------------- -----------------------
CHECKINTERVAL       5
CONF_FAIL           /mnt/flash/failed.conf
CONF_RECOVER        /mnt/flash/recover.conf
FAILCOUNT           2
HTTPTIMEOUT         10
IPv4                192.168.100.103
PASSWORD            4me2know
PROTOCOL            https
REGEX               Arista
TCPPORT             443
URLPATH             /eapi/
USERNAME            admin
VRF                 mgmt

Status:
Data                     Value
------------------------ -----------------------
CHECKINTERVAL:           5
CONF_FAIL:               /mnt/flash/failed.conf
CONF_RECOVER:            /mnt/flash/recover.conf
FAILCOUNT:               2
HTTPTIMEOUT:             10
HealthStatus:            UP
IPv4 Address List:       192.168.100.103
PASSWORD:                4me2know
PROTOCOL:                https
REGEX:                   Arista
Status:                  Administratively Up
TCPPORT:                 443
URLPATH:                 /eapi/
USERNAME:                admin
VRF:                     mgmt

Syslog Messages

May 30 16:48:45 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: %AGENT-6-INITIALIZED: Agent 'TCPCheck-TCPCheck' initialized; pid=12335
May 30 16:48:45 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: TCPCheck Version 3.0.0 Initialized
.
After HTTP Host goes down...
.
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Connection Timeout
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: HTTP HOST is down. Changing configuration.
May 30 16:49:30 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Applied Configuration changes from /mnt/flash/failed.conf
.
After Recover of HTTP Host...
.
May 30 16:49:57 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: HTTP host back up. Changing Configuration.
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_E: Enter configuration mode from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_I: Configured from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_E: Enter configuration mode from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 ConfigAgent: %SYS-5-CONFIG_I: Configured from console by root on UnknownTty (UnknownIpAddr)
May 30 16:49:57 DC1-SPINE-1 TCPCheck-ALERT-AGENT[12335]: Applied Configuration changes from /mnt/flash/recover.conf

INSTALLATION

Which method you use depends on your EOS release. The 4.26.0 boundary matters because that is where Sysdb mount profiles stopped being required, which is what the RPM existed to install.

EOS 4.26.0 and later - copy the script

No RPM, extension or mount profile is needed. Copy TCPCheck.py onto the switch, make it executable, and point the daemon at it. /mnt/flash is the easiest location because it persists across reloads.

Copy the file to the switch by whatever means you prefer, for example from the CLI:

copy scp://user@server/path/TCPCheck.py flash:

Then make it executable. The agent will not start without this, and the failure is not especially obvious in the logs:

bash sudo chmod +x /mnt/flash/TCPCheck.py

Confirm it took:

bash ls -l /mnt/flash/TCPCheck.py

You are looking for the x bits, e.g. -rwxr-xr-x.

Then configure the daemon:

configure
daemon TCPCheck
   exec /mnt/flash/TCPCheck.py
   option IPv4 value 10.1.1.1
   option PROTOCOL value https
   option TCPPORT value 443
   option REGEX value Arista
   option CONF_FAIL value /mnt/flash/failed.conf
   option CONF_RECOVER value /mnt/flash/recover.conf
   no shutdown

If you edited or transferred the file from Windows, make sure it has unix line endings before starting the daemon. See Agent crash due to Windows CR LF below.

EOS releases earlier than 4.26.0 - use the legacy RPM

Releases before 4.26.0 still require a Sysdb mount profile, which version 3.0.0 no longer ships. For those releases use the legacy RPM in the legacy/ directory of this repository:

legacy/TCPCheck-2.3.2-1.noarch.rpm

Install it as an extension in the normal way. It takes care of all the file requirements, including placing the mount profile in /usr/lib/SysdbMountProfiles and the agent in /usr/local/bin. With that RPM the daemon exec line is /usr/local/bin/TCPCheck, not /mnt/flash/TCPCheck.py.

Be aware of what you are getting: the legacy RPM is version 2.3.2, which is Python 2 only. It will not run on an EOS image that has removed the Python 2 interpreter, and it does not contain any of the fixes listed under What changed in 3.0.0. It is provided for continuity on older releases, not as a maintained alternative.

Sysdb mount profiles

Older releases of EOS required a SysdbMountProfile, and previous versions of this project shipped one as a second file named TCPCheck (no extension) alongside TCPCheck.py. That requirement also forced the agent filename and the mount profile filename to match.

From EOS 4.26.0 onward, mount profiles are no longer required. The SDK mounts what an agent needs automatically. Arista's guidance is to remove any mount profile you have, because running without one uses less memory and CPU than running with one.

As of version 3.0.0 this project no longer ships a mount profile. If you are upgrading from an earlier version, delete the stale file from the switch, since an RPM upgrade will not necessarily remove a file a previous version installed:

bash sudo rm -f /usr/lib/SysdbMountProfiles/TCPCheck

With the mount profile gone, the agent filename no longer has to match anything, so exec /mnt/flash/TCPCheck.py is fine.

If you are targeting an EOS release older than 4.26.0 you will still need a mount profile, and should use the legacy RPM in the legacy/ directory, described above.

Python 3 / EOS release support

Version 3.0.0 is a Python 3 port. It requires Python 3 and will not run under Python 2. Versions 2.3.2 and earlier require Python 2 and will not run on EOS releases that have removed the Python 2 interpreter.

TCPCheck version Python EOS releases How to install
2.3.2 and earlier Python 2 Pre-4.26.0. Tested on 4.20.1, 4.20.4, 4.20.5, 4.24.0. Requires a Sysdb mount profile. RPM in legacy/
3.0.0 Python 3 4.26.0 and later, where mount profiles are no longer used. Copy TCPCheck.py to /mnt/flash and chmod +x

The shebang is #!/usr/bin/env arista-python. arista-python is the EOS supplied wrapper that resolves to the correct interpreter for the image and sets up the eossdk import path, which avoids the agent silently changing interpreter across upgrades the way a bare python shebang does.

To confirm which interpreter the SDK bindings are installed under on a given image:

bash ls -d /usr/lib/python*/site-packages/eossdk*

What changed in 3.0.0

Changes required by Python 3:

  • socket.socket(_sock=...) has been removed. In Python 3 socket.fromfd() already returns a usable socket object, so the re-wrap that followed VrfMgr.socket_at() is gone. The explicit os.close() of the descriptor returned by socket_at() is still required, because fromfd() duplicates it.
  • ssl.wrap_socket() was deprecated in Python 3.7 and removed in 3.12, and ssl.PROTOCOL_TLSv1 went with it. HTTPS now uses an ssl.SSLContext. Certificate verification remains disabled, matching the previous behavior, so self-signed certificates still work. A side effect is that TLS is no longer pinned to TLSv1, so HTTPS checks now succeed against servers that have dropped it - which current EOS has.
  • base64.b64encode() takes bytes, and sockets send and receive bytes. The basic auth header, the request, and the response are encoded and decoded explicitly.

Defects fixed along the way, which were present in the Python 2 versions too:

  • The response is now read until complete rather than with a single recv(). A single read frequently returned only the HTTP headers, because servers commonly write headers and body as separate segments, which meant a REGEX intended to match page content could fail against a perfectly healthy server. The read uses Content-Length where present, so it does not wait for the server to close the connection.
  • A leading CRLF is no longer sent before the request line. RFC 9112 removed the allowance for servers to ignore an empty line before the request, and hardened HTTP stacks now reject it outright.
  • Socket cleanup moved into a finally block. Previously cleanup only ran on the success path, so an exception between connect() and cleanup leaked a file descriptor once per CHECKINTERVAL.
  • The error handler guarding socket creation called close() on the socket whose creation had just failed, raising UnboundLocalError and masking the real error.
  • CONF_FAIL and CONF_RECOVER file reads are guarded, and blank and comment lines are filtered before being passed to eAPI. Previously a trailing newline or a comment would cause the entire config batch to be rejected, and an empty file raised IndexError.

Every changed region in TCPCheck.py carries a comment explaining the reason, tagged PY3 PORT: where the change was forced by the language or standard library, and BUGFIX: where a latent defect was corrected.

Based on licensing, this is open source and this and other open-source tools are not supported directly by Arista Support. Support is best effort as it relates to extensions such as this one.

Troubleshooting

Agent reports down but the host is clearly up

Almost always a REGEX that cannot match. See Choosing a REGEX above. Confirm with:

bash curl -i <protocol>://<IPv4>:<TCPPORT><URLPATH>

Check three things in that output, in order: the status line is a 200 and not a 3xx; the body is the page you expect and not a redirect stub or a JavaScript shell; and your REGEX string is literally present in the body.

Note that curl follows redirects with -L while this agent does not, so a curl that succeeds without -i can easily disagree with the agent. Always use -i when comparing.

Agent crash due to Windows CR LF

If the agent is continuously crashing it might be caused by invalid characters in the python script. In the syslog outputs or show logging similar errors could be seen:

Aug 31 02:46:41 switch1 ProcMgr-worker: %PROCMGR-6-PROCESS_TERMINATED: 'TCPCheck' (PID=7319, status=32512) has terminated.
Aug 31 02:46:41 switch1 ProcMgr-worker: %PROCMGR-3-PROCESS_DELAYRESTART: 'TCPCheck' (PID=7319) restarted too often! Delaying restart for 120.0

To troubleshoot further the agent logs should be checked with bash cat /var/log/agents/<agentName>-<pid>.log (substitute <agentName>-<pid> with the actual name of the file

If the output is something similar as below, then it means the python script has an invalid \r which is the Windows carriage return (CR LF)

cat TCPCheck-Rack1-17880
==== Output from /mnt/flash/TCPCheck [] (PID=17880) started Sep 1 15:00:00.00000 ===
/usr/bin/env: 'python\r': No such file or directory

The solution is to convert the file to unix format, this can be done locally on EOS by editing the file with vi and typing :set ff=unix, so the steps would be:

  • drop down to global configuration mode and shutdown the daemon
       daemon TCPCheck
         shutdown
    
  • go to bash by typing bash
  • vi /mnt/flash/TCPCheck
  • type :set ff=unix
  • press Enter
  • press Esc
  • type :wq!
  • type exit to go back to EOS CLI and bring up the daemon again
  • no shutdown

Tip: When using Notepad++ to edit files always convert them to unix format by clicking on Edit - EOL Conversion and select Unix(LF) and save the file.

License

BSD-3, See LICENSE file

About

EOS SDK Script to test HTTP reachability and make switch config changes based on reachability

Resources

Stars

3 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages