Tuesday, September 16, 2014

AWS RDS MySQL SSL

Following my previous post of SSL with ELB's and instances I'd like to write a quick post about SSL and RDS MySQL. This is worth another blog post because it behaves slightly differently than your normal apache http or tomcat servers. 

This is because the SSL is enforced at the user level and not at the server port level. Typically when you disable insecure traffic to a particular server you will disable the port that it is listening on for insecure traffic - such as port 80 for the apache web server. 

MySQL is different, you still connect through the default port (3306), or whatever port you have configured it to run, but the difference is the way you connect. In order to ensure secure communication you must first create a user which requires SSL to connect:

First connect to the RDS instance as the root user:
mysql -h myinstance.123456789012.us-east-1.rds.amazonaws.com -P 3306 -u root -p
You can then create users in the specific way:
GRANT SELECT, INSERT, UPDATE, DELETE on db_name.* to 'encrypted_user'@'%' IDENTIFIED BY 'suprsecret' REQUIRE SSL;FLUSH PRIVILEGES;
You can now test this with the mysql client that you used earlier to connect to the database earlier:
mysql -h myinstance.123456789012.us-east-1.rds.amazonaws.com -u encrypted_user -p --ssl_ca=mysql-ssl-ca-cert.pem --ssl-verify-server-cert
The above should prompt you for the password that you set up above, and allow you to connect to the database securely. 

This is the database server end complete. The next part is to configure your application to use SSL to connect to the database securely. There are a number of ways in which you can complete this, for this example I'm going to configure a java application.

For this example we're going to configure the JDK to allow the secure connection. There are a number of options here, including application container specific configurations but this way has the advantage that all java applications (container or otherwise) will be able to connect.

First get the RDS MySQL server certificate from AWS:
wget https://rds.amazonaws.com/doc/mysql-ssl-ca-cert.pem
Now import this into the java trust store (replace $JAVA_HOME as necessary):
$JAVA_HOME/bin/keytool -importcert -alias rds -keystore $JAVA_HOME/jre/lib/security/cacerts -storepass changeme -file ./auspost-root-ca.cer -noprompt
Connecting to the MySQL database is the same as before however you now specify three new options in your connection string which connect using SSL:
jdbc:mysql://myinstance.123456789012.us-east-1.rds.amazonaws.com:3306?useSSL=true&verifyServerCertificate=true&requireSSL=true
For brevity I will not include the rest of the code that you use in java in order to connect.

Thursday, September 4, 2014

AWS and SSL - ELB to instance, Instance to ELB

At a glance

Here's the top tips to remember when attempting to enable SSL from an ELB to an instance and from that instance to another ELB:


  • ELB's don't like backend self signed intermediate/root certificates, they need the actual certificate that the instance server is presenting
  • Instances are quite happy to use the root certificate when connecting to a server presenting a signed certificate.
  • Openssl and curl are your friends for testing (explained in detail below):

openssl s_client -connect <elb_dns>:443 -CAfile server.crt

  • You can redirect your ELB listeners to listen on an un-secure port but communicate to the instance securely which is handy for testing
  • Use ELB health checks with SSL in order to get continuous verification that SSL is working


In the detail - What we're attempting to achieve


SSL (TLS) is the standard for encrypting traffic between a client and a server. One of the best explanations that I've seen about it is here so I'll leave you to read through (I know I found it useful) and get an overview of SSL. 

In this exercise we'll be attempting to enable encrypted traffic from an ELB to an instance and from that instance to another ELB. For example say if I have a web tier and an application tier, I would typically have an ELB in front of each, therefore I need to configure both front and back end certificates on each ELB and also the instances.

I'll be using a certificate which has been signed by an internal root CA, as it's a little more complicated and more like a scenario that you'll be presented with when trying to complete SSL with your own server. Openssl have docs on creating signing requests and self signed certificates.


ELB configuration

A great guide for creating HTTPS ELB's is already defined by the kind people at AWS so I'll only concentrate on the parts that tripped me up.

Front End

Front end certificates in AWS are stored in IAM, allowing you to choose them when creating new ELB's. This is pretty smart as you'll typically use the same certificates for many ELB's. Here, it's just a matter of selecting an existing certificate or uploading a new one per the guide.


Back End

This was a real pain. I assumed that ELB's would work in the same way that the bundle file would in your OS, that is, that you could upload the root certificate to the backend of the ELB and it would be enough to work with the certificate being presented with by the instance. However this was not the case, so my number one tip, is to forego trying the certificate chain and just upload the certificate that is being presented by the instance, uploading the certificate to the backend is describe in the the AWS guide.


Testing

Finally we get to do some testing. I completed this using openssl to check the front end certificate and then curl to test getting a page from my web server once it was configured (see below).

openssl s_client -connect <elb_dns_name>:443 -CAfile /tmp/root.crt
This will connect to your ELB and verify that the root certificate you have specified locally will work with the certificate being presented by your ELB. You should get an OK if everything is configured correctly.


Instance Configuration

In this case I'm going to go with a simple example and use an apache web server configuration for SSL. There are numerous detailed posts out there about configuring apache for SSL so I'll go over the very minimum that will pertain to the AWS ELB specifics. 


Front End

Your instance (in this case apache) will need to be configured to present a SSL certificate to the client (the ELB). In order to do this edit the httpd.conf or vhost conf (more info here) that you have configured to listen on 443 to have these values:

SSLEngine on
SSLCertificateFile /path/to/www.example.com.crt
SSLCertificateKeyFile /path/to/www.example.com.key

As we learned in the explanation of SSL, the certificate, which is just a fancy public key, only works correctly with a corresponding private key file in order to decrypt the data that the client is sending to the server. After restarting the apache server you should be able to view the certificates that the server is presenting:


openssl s_client -showcerts -connect localhost:443


Back End

This is where the configuration is a bit easier and you can generally use a root certificate as opposed to the certificate that is being presented by the ELB. The advantage of this is that the root certificate will generally be valid for longer meaning you don't have to change them as often. 

For this example I'm importing the root certificate to the operating system bundle file that most tools use by default, including curl. Working with RHEL 6.5 we can complete the following command to import the certificate:

openssl x509 -text -in /tmp/server.crt >> /etc/pki/tls/certs/ca-bundle.crt

This can then be verified by the following command:

openssl verify -CAfile /etc/pki/tls/certs/ca-bundle.crt /tmp/server.crt

Testing

Front end testing of the apache web server can be completed by redirecting your un-secure (port 80) listener on your web ELB to talk to the secure port of your instance, like the following:


You can then connect to you ELB via http and know that it's connecting to your instance securely. This is just an option when you want to test individual configurations.

For backend, not only will you want to verify that the certificate is installed correctly (using openssl verify), but you'll also want to check that using the updated bundle to connect to the ELB will work correctly, this can be done with the following:

openssl s_client -connect <elb_dns_name>:443 

This should return you a bunch of text (including the certificate presented by the ELB) but at the bottom you should get an OK if everything is configured correctly. 

Finally: End to End testing

Considering that you'll have a page on your web server (like an index.html file or something) you can use curl to get the file securely through the ELB. I completed this on a VM that I had already installed the root certificate in the ca-bundle with the following command:

curl -v https://<elb_dns_name>/index.html

If everything is configured correctly then you should not notice anything different that using simple http, and you should get your page displayed on the command line.

Added Extras

Constant SSL certificate Health Check

One tip that I got from a colleague of mine is to configure your health check on the AWS ELB to use SSL in the health check, meaning that you'll get continuous verification that the SSL certificate is valid between your ELB and instance. This is particularly useful when considering that certificates expire:


Verify ssl key matches certificate

Another handy thing to verify, if you're not generating the public/private key yourself, is to make sure that the certificate matches the key. This is explained very well in another article. 

Friday, January 4, 2013

Redirecting everything past root in apache


Ever needed to do some maintenance on an Apache web server? Well I do, and I had been wanting to redirect everyone to a maintenance webpage so that they would see that the temporary link instead of any of the actual website. 

I used to do this by putting in place a different configuration file for the Apache web server and allowing it to redirect to the maintenance page, with the following being the main line in the configuration file:

DocumentRoot /var/www/html/maintenance

This was a pain for two reasons:

  1. It meant that I had to go in and save the actual config file off to another temp file, then rename the maintenance config file to be the actual config file
  2. When anyone requested a webpage that wasn't the DocumentRoot of the website it would not get redirected to the maintenance page

So after some ooogooglygoogling I was able to find the following nugget:

RewriteEngine Off
RewriteCond   %{REQUESR_URI} !=/index.html
RewriteRule   ^ /var/www/html/maintenance/index.html

in the configuration file of the actual website Apache config. This means that I can keep the same configuration file for my maintenance and my actual site, doing away with any pesky moving and renaming etc. and switch between the maintenance and production with just putting the rewrite engine off or on.


It also means that anyone requesting a webpage that is not at the route of the website will be directed back up to the maintenance page. It does this by matching anything from the start of the line - denoted by the '^' i.e. the start of the line, and redirecting to the html file denoted in the second part of the RewriteRule directive.

This is much easier for me maintaining less files, and a more natural feel when people have to be redirected going to the site.

Monday, October 29, 2012

First attempt at some sort of continuous deployment

So I've heard of this continuous deployment concept and it all sounds great. I am keen to give it a go myself but am unsure where to start. So rather than attempt this with a production application in work I pick a nice small app that I've written to manage some of our application config.

This is a small web application hosted on one of our servers which has a database for persistence. Not much you might say but still has got the ingredients to try some sort of automatic deployment. To give you an idea of what I was trying to get away from I need to describe the evolution of the deployment.

What I started with

When writing the application I was not at all concerned about deployment, I could barely write html never mind get the thing running on an actual server. However as time went on I realised that I would have to have a go at getting this deployed out to a running server. This started with a very very manual process. First I had to install all the prerequisites, in this case python 2.7, mysql server, mysql-python connector, django and all those other little good things that it needed to run.

After that it's a matter of getting the code on the machine. For this I had a bunch of hand cranked scripts which replied heavily on ssh keys on my local machine to do the deployment. Not ideal. This really became apparent when other developers wanted to get in to make some changes. I wanted to remove this reliance on me as a bottle neck for deployments.

Conversion to rpm's

In step rpm's, Red Hats package management system. This was something we'd been trying to investigate in work of how to package our production code so I wanted to see how I could implement this as part of a trail run. This took care of the any scripts that had to run post deployment as all is catered for in the rpm word. A better solution alright however I was still taking the rpm down to my local VM after CI had built it and doing manual testing after an export from the live database.

Automated deployment to a test environment

In order to take me and my lovely VM out of the picture I really needed a test environment where I could tear down the database and recreate as necessary. This meant that I would be able to automatically deploy to this environment when the CI build had completed. A much better solution. In order to do this I had to get a new environment up and running, back to compiling and install python again. To hell with this I said - in order to be able to do this anywhere I really needed packages for all my dependencies as well.

I was much more confident in rpm's now so I created one for all the dependencies also, I now have a python rpm, a mysql-python rpm, a mod_wsgi rpm. All ready to set up on a new environment if I so wish. There was another angle on this as well - any developer could get up and running with a developer environment in no time, assuming they would use Red Hat or one of it's alternatives (CentOS in my case)

Current State

I'm now able to deploy to the test environment my new rpm from the build. I also scripted an export of the database, a fresh if you will before the deployment so that the test database would be as close as possible to live data. This means now more testing on my local VM to make sure that the rpm and deployment is up to scratch.

Next Steps

What now. Testing that's what. And lots of it. I'm of the opinion that in order for me to get to the point where I'm able to click the button to deploy to live, i.e. that I've deployed to the test environment and am able to test the change I need a bunch of integration tests. Something that I didn't do well in the first line of development, yes to my shame I cut some corners on the testing front.

I'm really feeling the pain of this architectural blunder now. I have very little confidence in what I add into the app will not regress some other part of it. The technical debt that I have to pay down is quite substantial on this front. There was a second side affect to this I didn't think of either. A social get out clause that other developers were able to see that I hadn't put effort into proper testing and were able to not bother with it themselves, further increasing the technical debt within the application.

Definitely the lessons learned here are that testing, while not an immediate concern at the time of writing the application, is key to be able to take the human out of the picture. I now have a deeper respect when jumping into writing a part of an application to the tests that need to be in place to support the change that I'm making and the resultant lack of confidence in my change if this is not done at the start.

Testing through your home firewall router

Recently I came across a problem where I needed to test a hosted service getting access to a port on my laptop. The general workflow was that I would initiate the request from the hosted service which would then talk to a service running on a tomcat on my local laptop.

There are a number of problems to get around when attempting to do this. First you can't just use your IP address as you see it on a command prompt ipconfig or ifconfig. The network address that you see here is the internal address that your home router has given you and any device connected to it. Therefore the hosted service isn't going to have a clue how to get to 192.168.0.*

What you need here is the external IP address from your home router. There are a number of ways to get this, including some websites that will display it for you but I went to the source - the router. Luckily there was a configuration page which told me the external IP that my router was using.

The next problem is your firewall. Most modern routers come with a firewall - for good reason - to stop those nasties from getting in to your local network. For this test I was testing for a short time so I was happy to bore a hole through my firewall to my local computer IP through port forwarding. This is where you select a port on the router and a port on your laptop and any traffic hitting your router via it's external IP will be forwarded to that port on your laptop through it's local network IP.

In my case I just forwarded port 80 to port 80 (the default) as I was able to set up my service on my local laptop on that port. So after all this was done I was ready to test. Initiating the test from the hosted service still did not work however...ragin.

The final thing that you need to complete is to turn off your windows firewall (if you're running windows). Since I had already let the traffic through the router it was the windows firewall that was blocking the traffic.

Very happy with myself for actually getting that done as I've never tried it before. Good fun.

Oh and yes, I did quickly turn on my firewall again after the testing was complete and remove the port forwarding!

Wednesday, June 13, 2012

MySQL GRANT oddities

The problem

Just yesterday after setting up a mysql slave I tried to switch the backups of the master to the slave, trying to reduce the load on the master. While trying the following command:


mysql> GRANT ALL ON *.* TO 'backupuser'@'localhost';
I got the following error:

 Access denied for user 'root'@'localhost' (using password: NO)
hmmm, quite curious since I was logged into the mysql server as root. After some googling on the subject I found a weird little problem with the tables between versions.

This error above only happened because I had upgraded the mysql server from a 5.0 to a 5.5 server. I had thought there wasn't much to this, I also changed the location of the data directory to be a new partition to handle the amount of data that I had.

The crux of the issue was that during the upgrade I moved the data directory contents as well. Meaning that the old tables were in place with the new server.

The Solution

What I didn't realize was that I needed to upgrade the mysql tables from the old version to the new. So with the following command:

shell> mysql_upgrade
I got the above error again. 

 Access denied for user 'root'@'localhost' (using password: NO)
Damn! But not to worry, with the above command is able to take  a username and password:

shell> mysql_upgrade -u root -p
Nice one I'm now in and it gives me a bunch of errors. Damn! Error code 13? Hold on that's a permissions error. I wonder if running it as the mysql user that has ownership of the data file tables will work:

shell> sudo su - mysql
shell> mysql_upgrade -u root -p
Nice, worked a charm this time.

I can now go back to my little grant that I was trying to do in the first place:

mysql> GRANT ALL ON *.* TO 'backupuser'@'localhost';
And I'm able to complete it. Nice one.

Friday, June 8, 2012

Implementing Kanban

Frustrations


A general frustration of mine of using an agile type board to track the work the team I am currently working in was that we had very little insight into what types of work that we completed. I had worked in a development team a while back and was very impressed about how good it can work for a team on a specific project. About a year and a half ago I moved into the DevOps type role, still a software developer by title but more concerned about the build and release side of things. The board that we implemented for this was the following:






A couple of things to note about this board. 

  • There's no rhyme or reason for the colors that are used
  • The backlog is huge and rarely moves as things get put directly into next
  • All the backlog items are "planned" work things that we need to get to at some point but very rarely get the time for

Along with the changing role there were a great deal of work type baggage. For instance we might have to change build scripts, promote code, build the latest release candidate, fix a problem with an environment. The main change was that a good portion of this work was not planned, all very ad hoc. People coming up to our desk and saying "I need a build" or "the MQ server is down", stuff like that.


The problem with the above board is that none of this ad hoc work is visualized. This meant that the planned work sat in the backlog for a very very long time and the in progress stories were there for a long time too. Things would also come straight into next without be prioritized in the backlog either. So all in all, although being very busy, it felt like we were doing nothing cause the board wasn't moving.


Trying something different


I was then lucky enough to get on a conference about Lean. A couple of speakers at it specifically talked about Kanban for DevOps. This is a growing subject and I was very interested in how they applied lean/kanban to there teams work. It talked alot about visualizing the types of work that are completed by the team, showing the dependencies on other teams and how to make changes to the current way of things.


From this then we created a new board, after the first couple of days it looked something like this:




The things to note are:

  • There are three main swim lanes one for each type of work that we think we do, incidents (yellow), change requests (blue) and planned work (green) with a waiting for or impediment (pink) for problems
  • We have a lot less in our backlog
  • Blockers and waiting for are visualised

The new swim lanes go a long way to showing the actual types of work that are coming across our desk. As you can see the different colors show that a good proportion of the work is actually incidents. These are not typically large jobs however they do involve a context switch which is time consuming in it's own right.


The difference it makes


There are a couple of differences we say this could make to the team.
  • Allows people coming up to the desk to see what work is currently going on and how there task would fit into the current work load
  • Shows the flow across the board, meaning that it actually looks like we're doing something - this is more personal for me as I was able to feel better about the work getting done
  • Data - it might not be apparent from the picture but we track how long something is sitting in the board with a start date and how long it is in progress with a dot for every day


Next Steps


Next steps will be what to do with the data that is being captured from the board. We are planning to clear the board every week and capture the data that we can get and then make some improvements. Cause that's the end game really, is to make some genuine improvements into the way that we work. This will mean that the "done" column will also be cleared down and the board will look fresh for the new week coming.


I'm still undecided whether or not there is enough data there to be useful to us, and what we're marking down is going to tell us anything. But what we record will possible change by having a look at what is useful to us with the data that we currently have.


Anyway, I might post something in a month or two about this to see if the same holds true, if we've made any improvements or just clean scrapped the board for something else.