* Take all data that caused a failure in the application and have a script that generates that in the test database. Maybe it is a unicode character that was misinterpreted etc. This makes a basic regression test easier to apply.
* Take some percentage of the Production database and have a script or application obscure the data and load it into the test environment. Generally I like to make this configurable so it can be used to generate a small dataset on a devs laptop, or a larger dataset on a test server.
* Obscure the data differently every time the application runs, and have the data stay within the same range as the original production data. For example, if a currency field showed $199.00, then you might obscure it to be $179.00, use a formula to do this. If a string contained 4 spaces, 2 capitals and 3 punctuation marks keep the ratios the same just obscure the character values (I do not obscure punctuation for testing).
I point out the obscuring because it is the easiest way to help prevent client data from being lost, stolen or abused. I like to write an application that copies the data over and lets you pick how much data to use. For example, for a dev laptop, maybe they can only get 1% or 10k complete records. For testing server it can get anything from 20% to 100%. Also, if you do it this way, you can have a developer database generated every day overnight so it is ready for anyone who needs to just copy it local in the morning for test, and for the test server you can regenerate it on some rolling basis to keep it fresh. Also, have the application reseed the cache from the newly generated test database.
Are you also trying to seed your production environment?