The Wayback Machine - https://web.archive.org/web/20160312233229/http://planetpython.org/

skip to navigation
skip to content

Planet Python

Last update: March 12, 2016 09:50 PM

March 12, 2016


Podcast.__init__

Episode 48 - PyData London with Ian Ozsvald and Emlyn Clay

Visit our site to listen to past episodes, support the show, join our community, and sign up for our mailing list.

Summary

Ian Ozsvald and Emlyn Clay are co-chairs of the London chapter of the PyData organization. In this episode we talked to them about their experience managing the PyData conference and meetup, what the PyData organization does, and their thoughts on using Python for data analytics in their work.

Brief Introduction

Linode Sponsor Banner

Use the promo code podcastinit20 to get a $20 credit when you sign up!

Hired Logo

On Hired software engineers & designers can get 5+ interview requests in a week and each offer has salary and equity upfront. With full time and contract opportunities available, users can view the offers and accept or reject them before talking to any company. Work with over 2,500 companies from startups to large public companies hailing from 12 major tech hubs in North America and Europe. Hired is totally free for users and If you get a job you’ll get a $2,000 “thank you” bonus. If you use our special link to signup, then that bonus will double to $4,000 when you accept a job. If you’re not looking for a job but know someone who is, you can refer them to Hired and get a $1,337 bonus when they accept a job.

Interview

Keep In Touch

Picks

The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

Visit our site to listen to past episodes, support the show, join our community, and sign up for our mailing list.Summary Ian Ozsvald and Emlyn Clay are co-chairs of the London chapter of the PyData organization. In this episode we talked to them about their experience managing the PyData conference and meetup, what the PyData organization does, and their thoughts on using Python for data analytics in their work.Brief IntroductionHello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.Subscribe on iTunes, Stitcher, TuneIn or RSSFollow us on Twitter or Google+Give us feedback! Leave a review on iTunes, Tweet to us, send us an email or leave us a message on Google+Join our community! Visit discourse.pythonpodcast.com for your opportunity to find out about upcoming guests, suggest questions, and propose show ideas.I would like to thank everyone who has donated to the show. Your contributions help us make the show sustainable. For details on how to support the show you can visit our site at pythonpodcast.comLinode is sponsoring us this week. Check them out at linode.com/podcastinit and get a $20 credit to try out their fast and reliable Linux virtual servers for your next projectI would also like to thank Hired, a job marketplace for developers and designers, for sponsoring this episode of Podcast.__init__. Use the link hired.com/podcastinit to double your signing bonus.Your hosts as usual are Tobias Macey and Chris PattiToday we are interviewing Ian Ozsvald and Emlyn Clay about their work with PyData London, a group within the PyData organization. PyData London represents the largest Python group in London at ~2850 members, they hold regular monthly meetups for ~200 members at AHL near Bank and a yearly conference for around ~300 members. Last year, they and their sponsors raised over £26,000 to sponsor the development of core numerical libraries in Python. Use the promo code podcastinit20 to get a $20 credit when you sign up! On Hired software engineers designers can get 5+ interview requests in a week and each offer has salary and equity upfront. With full time and contract opportunities available, users can view the offers and accept or reject them before talking to any company. Work with over 2,500 companies from startups to large public companies hailing from 12 major tech hubs in North America and Europe. Hired is totally free for users and If you get a job you’ll get a $2,000 “thank you” bonus. If you use our special link to signup, then that bonus will double to $4,000 when you accept a job. If you’re not looking for a job but know someone who is, you can refer them to Hired and get a $1,337 bonus when they accept a job.InterviewIntroductionsHow did you get introduced to Python? - ChrisWhat is the PyData organization, how does PyData London fit into it and what is your relationship with it? - TobiasIn what ways does a PyData conference differ from a PyCon? - TobiasDoes PyData do anything in particular to encourage users from disciplines that might not be aware of how much our community has to offer to choose the Python suite of data analysis tools? - ChrisYou have both spent a good portion of your careers using Python for working with and analyzing data from various domains. How has that experience evolved over the past several years as newer tools have become available? - TobiasFor someone who is just getting started in the data analytics space, what advice can you give? - TobiasHow can conferences like PyData help strengthen the bonds and synergies between the Python software community and the sciences? - ChrisThere are a number of different subtopics within the blanket categorization of data science. Is it difficult to balance the subject matter in PyData conferences and meetups to keep members of the audience from being alienated? - TobiasData science is a young field and we've yet to see lots of examples of the successful use of data. How are London-based compa

March 12, 2016 04:01 PM


Will McGugan

Progressive loading of high DPI photos

I spent an evening adding 'progressive' loading of the title images to my blog.

The title images for this blog are 3840 × 2160 and a hefty ~650K each. That's entirely intentional; as a photographer I wanted them to look as sharp as possible and take advantage of high pixel density screens.

The only downside of hi-res photos is that even with a good internet connection, you can still see the photo loading as the browser reads the JPEG. It's visually jarring and way too reminiscent of the web, circa 2000s.

A reasonable solution is to first download a smaller lower-resolution version, then load the full resolution image on top of that. So the user sees something relatively quickly, without the visual contrast of an image loading on a blank background.

JPEG images natively support something very similar; if you save your photo as a progressive JPEG, you will get a pixelated image that gradually gets more detailed as the browser reads the file. Personally, I don't like the visual effect of progressive JPEGs. They may be an improvement over displaying a line at a time, but the initial pixelation is not visual pleasing.

This blog implements an alternative method of progressive loading of photos. For each title photo there is a 1920 × 1080 version that is only around 70K or so. This smaller version of the photo is heavily blurred with a gaussian blur filter by Moya Tech Blog. The blurred image gives a good impression of what the full image will look like, without obvious pixelation (and conveniently blurred photos compress really well in comparison to the non-blurred original).

Here's an example of one of the blurred preview photos:

Image

Blurred preview of a hi-res photo

This blurred photo is absolutely positioned directly behind the title image, so that when the hi-res version loads it covers the blurred image. There is also a CSS transition to smoothly fade in the hi-res version, further reducing any visual interruption.

All in all, it works quite well. Technically, it will increase page-load time since it serves an extra image, but from the visitor's point of view, the page will appear to load significantly quicker.

If you are reading this via a feed, you might want to go directly to my blog and click around the links at the top to see the effect in action.

March 12, 2016 02:43 PM


Andre Roberge

Multilingual Reeborg

Instead of separate versions for different languages, (English, French and Korean so far), Reeborg's World is now available in all languages at a single location.  Mixed language modes are supported (e.g. UI in Korean with Python programs using English commands like move() ).  This previous single-language versions are still available (English, French, Korean) and will remain so until the new

March 12, 2016 09:46 AM


Import Python

ImportPython Issue 64


Word From Our Sponsor

Image
Python Programmers let companies apply to you, not the other way around. Receive interview offers which include salary and equity information. Companies see each other's offers, and compete for your attention. Engage only with the companies you like. REGISTER

Worthy Read

django
Channels is an exciting upcoming feature of Django that will allow Django sites to support use cases that usually required the use of external tools and libraries (even non-Python ones) and even has the potential the way we work with the framework entirely.

machine learning
K-Means Clustering is a machine learning technique for classifying data. It’s best explained with a simple example. Below is some (fictitious) data comparing elephants and penguins. We’ve plotted 20 animals, and each one is represented by a (weight, height) coordinate.

python3
Since its debut in 2008, Python 3 has come a long way. Gone are the days when it lacked support for almost all useful libraries and tools. Python 3 offers many improvements and amazing new features that make writing robust code in Python easier than ever. In this article, Toptal engineer Dario Bertini discusses some of the improvements and features that Python 3 has to offer, and explains whether switching to Python 3 is a smart choice right now.

django
A Django implementation of JSON Web Token Authentication (https://tools.ietf.org/html/draft-ietf-oauth-json-web-token-32).

django
,
docker
These diagrams shows the Docker images needed to set the application service in his first version (App:1.0). The version number will be increasing during the system evolution.

pandas
,
json
Working with large JSON datasets can be a pain, particularly when they are too large to fit into memory. In cases like this, a combination of command line tools and Python can make for an efficient way to explore and analyze the data. In this post, we’ll look at how to leverage tools like Pandas to explore and map out police activity in Montgomery County, Maryland. We’ll start with a look at the JSON data, then segue into exploration and analysis.

pycon
At this moment, exactly 2,000 people are registered for PyCon 2016 — which puts us ? of the way to capacity!. If you are planning to visit now's the time to book the tickets.

core python
American Fuzzy Lop is both a really cool tool for fuzzing programs and an adorable breed of bunny. In this post I'm going to show you how to get the the tool (rather than the rabbit) up and running and find some crashes in the cPython interpreter.

Recently we came to a conclusion that it's enough. We agreed that the backend of the service is going to be first in line. Because all of it was about to be rewritten we thought that maybe it's a perfect time to evaluate other technologies. Currently it's using Python2.7 + Django + Gunicorn. We consider going to either Node.js with Express 4, Python3 with aiohttp or C++.

What is the best Python book for experienced programmers? My background is in Ruby, C++, JavaScript (and a little Clojure) . I told a prospective employer that I knew Python. So I need to know Python.

interview
This week we welcome Chris Moffitt (@chris1610) as our PyDev of the Week! Chris has been an active writer about Python on his blog and a speaker at DjangoCon.



Projects

neural-doodle - 367 Stars, 13 Fork
Turning your two-bit doodles into fine artworks!

SSHKeyDistribut0r - 97 Stars, 7 Fork
A tool to automate key distribution with user authorization

resume - 33 Stars, 2 Fork
Automatically generate your résumé and various cover letters from YAML files.

ascii_qgis - 9 Stars, 3 Fork
A ASCII QGIS map viewer

pyphoon - 8 Stars, 0 Fork
ASCII Art Phase of the Moon (Python version)

AnimeWatch - 6 Stars, 1 Fork
Front End for mplayer and mpv

shahinday - 6 Stars, 0 Fork
A terminal based game written in Python. Extend it and contribute! Put together very quickly, no comments, but kind-of self-documenting...

slackStocks - 5 Stars, 1 Fork
Slackbot for stock prices

packtSnatch - 4 Stars, 0 Fork
Script for ordering and downloading ebooks from packtpub freelearning page

March 12, 2016 01:40 AM

March 11, 2016


Weekly Python StackOverflow Report

(x) stackoverflow python report

These are the ten most rated questions at Stack Overflow last week.
Between brackets: [question score / answers count]
Build date: 2016-03-11 17:32:31 GMT


  1. sine calculation orders of magnitude slower than cosine - [25/2]
  2. Python eval: is it still dangerous if I disable builtins and attribute access? - [17/4]
  3. Cleanest way to obtain the numeric prefix of a string - [16/7]
  4. Python 3: super() raises TypeError unexpectedly - [14/2]
  5. `object in list` behaves different from `object in dict`? - [13/2]
  6. how to print 3x3 array in python? - [12/7]
  7. Is there a one line code to find maximal value in a matrix? - [9/4]
  8. Unpacking arguments from argparse - [9/3]
  9. Python Iterate through list of list to make a new list in index sequence - [7/4]
  10. What's the best way to "periodically" replace characters in a string in Python? - [7/3]

March 11, 2016 05:45 PM


Investing using Python

Forecasting anything using anything with naive Bayes, Python, pandas (sample strategy)

This is huge subject, so I'll try it cover very fast. I'll try to find relationship between some data (in this case signal would be XLF, financial sector ETF, delayed by 1 day) and target would be S&P500 futures in the CFD form. Probably we'll go long on zero upside probability. Logically zero probability doesn't make any […]

March 11, 2016 03:25 PM


PyCharm

PyCharm 5.1 Beta 2 is available

Today we announce the release of PyCharm 5.1 Beta 2, an updated feature-complete preview version of the future PyCharm 2016.1. The build #145.256.43 is already available for download and evaluation from the EAP page.

Yesterday we announced the new versioning model for all JetBrains Toolbox products. With the nearest release, former PyCharm 5.1 will become PyCharm 2016.1. The new versioning will be aligned with releases of other JetBrains tools, following the new YYYY.R format. Please read more about new versioning in the JetBrains Toolbox—Release and Versioning Changes blog post.

Important note: As the versioning migration process takes some time, currently build #145.256.43 has old versioning (PyCharm 5.1 Beta) inside the build. We’re going to finally change versioning with the PyCharm 2016.1 Release Candidate which is planned for next week.

The full list of fixes and improvements for this build can be found in the release notes.

Download PyСharm 5.1 Beta 2 for your platform and please report any bugs and feature requests to our Issue Tracker. It also will be available shortly as a patch update from within the IDE (from the previous EAP build only) for those who selected the EAP or Beta Releases channels in the update settings.

Stay tuned for a PyCharm 2016.1 release announcement and follow us on Twitter.

The Drive to Develop
-PyCharm Team

Image

March 11, 2016 01:39 PM


Jamal Moir

How to Use Google's Python Client Library to Authorise Your Desktop Application With OAuth 2.0

Image
In a previous post I mentioned how I'm going to stop using Blogger's built in post editor due to the horrendous HTML is produces. Well, I had no luck finding a desktop blogging client that worked well. The existing blogging clients either don't work on linux or development was stopped some ten years ago.

As such, I am now developing my own desktop blogging client in Python. You can view the project on GitHub.

One of the things that was a bit a pain to figure out was authenticating the desktop client with Google's API using OAuth 2.0. Personally I don't think it's very well explained on Google's website and find the site uncomfortable to navigate. So, for the convenience of all of you that want to connect to and authorise a desktop application with Google's API, here's how to do it.

GETTING THE CLIENT LIBRARY

First off, we need to download the Python client library. I'm going to assume that everyone reading this blog is using pip, if you aren't... Start using it. If you are one of the elite using Python 3 (did I ruffle a few feathers?), then lucky you, you should already have in installed. If you don't have it, google is your friend.

In our terminal we execute the following command:

$ pip install --upgrade google-api-python-client
...And that's it, well done.

CREATING A NEW GOOGLE APIS CONSOLE PROJECT AND DOWNLOADING YOUR CLIENT SECRET

To use Google's API, we need to have a google account (you have one, right?) for accessing the Google Developer's Console with. 

We then navigate to the Developer Console's projects page and create a new project for our application by clicking the 'Create project' button and filling in the form that pops up.


Image
Enter your projects name and hit create.

Then we get redirected to our project's dashboard. On this dashboard there is a large blue box saying 'Use Google APIs' which we click.

Image
Click this to be taken to the Google APIs page.
We then get taken to a page which displays all of the APIs available to us; there are lot's and select the API that we are planning on using, I will be using blogger as an example.

Once we've selected the API we will be using, we will again be redirected to another page, on this page there will be a button that says 'Enable', clicking this lets us use the selected API. 

We are then presented with a warning box that prompts us to create credentials, which is exactly what we will do.

Image
Click the 'Go to Credentials' button.
We then get taken to a new page with a few options for us to fill in. We select the version of the API we want and for 'Where will you be calling the API from?' select 'Other UI (e.g. Windows, CLI tool)'. Also, we will be accessing user data so select the 'User data' option and click the 'What credentials do I need?' button.

Image
Fill in the appropriate details and hit the blue button.

Then go through the next two steps; 'Create an OAuth 2.0 client ID' and 'Set up the OAuth 2.0 consent screen' and input the information that applies to us. Download the credential information if you like and click the done button.

We get taken to a page listing credentials. Next to the credential we just created there is a download button, press that to download the 'client secret' which we will need later and move it to the root directory of your project.

Image
Download the client secret by clicking the circled button.

That's all we need to do with the Google Developer's Console. Next, onto the code.

USING GOOGLE'S PYTHON CLIENT LIBRARY TO AUTHORISE YOUR APPLICATION

The code consists of four steps:
  1. Getting an authorisation code
  2. Exchanging the authorisation code for credentials
  3. Creating an httplib2.Http object and authorising it using the credentials
  4. Creating an API service object to make calls to the API

THE CODE


def get_credentials():
"""Gets google api credentials, or generates new credentials
if they don't exist or are invalid."""
scope = 'https://www.googleapis.com/auth/blogger'

flow = oauth2client.client.flow_from_clientsecrets(
'client_secret.json', scope,
redirect_uri='urn:ietf:wg:oauth:2.0:oob')

storage = oauth2client.file.Storage('credentials.dat')
credentials = storage.get()

if not credentials or credentials.invalid:
auth_uri = flow.step1_get_authorize_url()
webbrowser.open(auth_uri)

auth_code = input('Enter the auth code: ')
credentials = flow.step2_exchange(auth_code)

storage.put(credentials)

return credentials

def get_service():
"""Returns an authorised blogger api service."""
credentials = self.get_credentials()
http = httplib2.Http()
http = credentials.authorize(http)
service = apiclient.discovery.build('blogger', 'v3', http=http)

return service


WHAT'S GOING ON

scope = 'https://www.googleapis.com/auth/blogger'

flow = oauth2client.client.flow_from_clientsecrets(
'client_secret.json', scope,
redirect_uri='urn:ietf:wg:oauth:2.0:oob')

First in get_credentials() we create a client object from the client_secret.json file that we downloaded earlier. We need to specify the scope, and the redirect_uri.

The scope declares the API and the level of access that we will be using, you can find a list of scopes here. The redirect_uri is how the response will be sent to our application, for more information on redirect uris read this.

storage = oauth2client.file.Storage('credentials.dat')
credentials = storage.get()

This code allows us to store the credentials so we don't have to re-authorise the client every time. It creates a Storage object and loads credentials.dat, which contains our credentials. If it doesn't exist it gets created. It then gets the credentials from the storage object.

if not credentials or credentials.invalid:
auth_uri = flow.step1_get_authorize_url()
webbrowser.open(auth_uri)

Then it checks to see if the credentials either don't exist or are for some reason invalid. If this check passes it means that new credentials are required so it generates an authorisation url and then opens it in the system's default browser.

auth_code = input('Enter the auth code: ')
credentials = flow.step2_exchange(auth_code)

storage.put(credentials)

On the page that is opened in the default browser, the user is presented with a code and then prompted to input it into the application. We then ask the user to enter the authorisation code that they were given and use it to generate credentials. Obviously there are better and more user-friendly ways to handle the inputting of the code. The credentials are then stored for later.

credentials = self.get_credentials()
http = httplib2.Http()
http = credentials.authorize(http)

Next, in get_service() we get our credentials using the get_credentials() function we made earlier and use it to authorise an httplib2.Http object which the client library will use to issue HTTP requests.

service = apiclient.discovery.build('blogger', 'v3', http=http)

Finally we create the API service object which we can use to interact with the API and we are finished.

Now your application will be able to authorise itself, connect to your api of choice and interact with it via Google's client library. It took me a while to navigate googles documentation and a little bit of trial and error to get this working properly, but hopefully this post will allow you to skip over that and get right into the actual building of your application.

If you liked this post or found it helpful, please share it so other people can find it too. Also, don't forget to subscribe to my posts feed so you don't miss anything.
Image

March 11, 2016 12:46 PM


Kushal Das

Python track at FOSSASIA 2016

Image

(Stephen J. Turnbull during his talk in Python track, FOSSASIA 2015).

The 2016 edition of FOSSASIA is happening from 18-20th March, in the Singapore Science Center, Singapore. We are again having a Python track in this event, which starts on 19th March (Saturday). The following is a summarised entry of the schedule for us.

On 19th March

On 20th March

If you look closely at the talks, we have talks on testing framework (pytest), talks on particular applications including how Python is being used to teach science in class rooms (ExpEyes). We have hands on workshops about Python3, Telegram, and Twitter API, and about symbolic computation. It is going to be lot of discussions, and learning. The full conference schedule is also live. It has many other tracks, and speakers from different upstream projects. We also have a DevOps track, which is filled with some excellent talks from my colleagues in Red Hat.

March 11, 2016 12:38 PM


Tim Golden

Planet Python: an unintended trip down Memory Lane

So… you may have noticed some rather antique posts popping up on Planet Python this morning. This is the result of two things: our clearing out the cache to try to address an issue [*]; and the way in which the feed parser tries to pick up the dates from the feeds it’s scanning.

For now, treat it as a unique chance to view the Python blogosphere through the spectacles of yesteryear! The older posts should just drop down as others are updated, but if they don’t we’ll do a little surgery to the config. In any case we’ll try to do some code housekeeping against any future outbreaks of nostalgia.

[*] … and which didn’t fix the problem!

March 11, 2016 11:30 AM


Malthe Borch

The Sign Up With Google Mistake You Can't Fix

Tue, 1 Mar 201618:30:00 GMT – Updated

This blog post has been revised to more accurately reflect some details. Thanks to Fleep CEO Henn Ruukel for reaching out.

Today – by regrettable oversight – I exported my recent e-mail history to Fleep, a collaboration platform used by a new client, and let their software synchronize future e-mails. I thought I was just signing up, giving up my basic info.

The mistake was all mine. I did not notice that the authentication screen was requesting for permission to allow Fleep to "manage my e-mail". But I did realize the consequences moments later when I saw that many of my contacts and my recent e-mail history had been pulled into their system.

It turns out that Fleep actually only pulls down the most recent 200 e-mails. But this wasn't apparent at all to me because I wasn't expected any of my e-mail to be imported. I had unwittingly allowed an app to download my entire personal correspondence for about a decade.

While I do blame Fleep for using what I perceive to be a euphemism – connect with gmail – and not making it crystal clear that you're not just signing up for a messenger app but actually fully integrating your e-mail account, the bigger problem here lies with how Google makes this possible:

It's just too easy to give away your personal information on the internet and this needs to be fixed.

We have a similar problem with apps on mobile devices that ask for permission to access all photos in order for you to be able to select just one. I think for the most part you can trust apps to do the right thing – but the way it's currently set up, there is no transparency in what these companies do with the data you have granted them access to.

Legally, I think the EU Data Protection Directive has me covered, but once you have handed over your data to an internet company, it's really out of your control.

Big internet companies, please take privacy seriously and help your users understand the consequences of their actions.

March 11, 2016 07:53 AM


SDJournal

Exhedra: a conferencing/forum application in Django

I've recently started some development on the Exhedra project (a conferencing/forum application) using Django. For anyone reading this post on SDJournal, the posts from this category are also syndicated on the Django Community page. Rather than clutter up that page with multiple posts about this project I've set up a separate weblog to cover it.

March 11, 2016 07:50 AM


David Marte

Weblog in Python and Google App Engine

A friend of mine rhetorically posed a question: How Difficult would it be to start a blog from scratch? Well, it won't really be from scratch. Nothing nowadays start from scratch. If kids of the 80s were the last to listen to cassette records and the first play with dvds, programmers of the 80s were the last to ever really write things from scratch, unless you are writing in PURE-C today. If you are like me, you call yourself a programmer even though you can't start a project if noone has yet written a convenience API or library.

March 11, 2016 07:50 AM


Juri Pakaste

Talking to servers: WebSockets

This is a sort of follow-up post to ZeroMQ, iOS and Python from a year ago. I again wrote a test app and server. Of the earlier components, iOS stayed, but ZeroMQ I replaced with WebSockets and the server is this time written in Clojure.

Idea of the exercise is the same as last time: two-way communication between an iOS client and server over a persistent connection. WebSockets is a more mainstream technology than ZeroMQ with a wide variety of servers available, even if it is a young spec and not quite everything has stabilized yet.

On the iOS side there's two options, short of writing the protocol implementation yourself: some kind of horrible Javascript-UIWebView bridge and Square's SocketRocket. I went for SocketRocket. It's available from CocoaPods and at least in this brief test worked fine.

Last time I was planning on doing the server in Clojure but ran out of patience. This time the technology stack was easier. I pretty much followed Jay Fields' example with just a few changes here and there and had a server running in no time.

The system doesn't do anything fancy, or do it particularly beautifully: the app waits for a button tap, sends a number wrapped in JSON to the server, the server increments it and sends it back.

The code is on BitBucket as always. Enjoy.

March 11, 2016 07:49 AM


Lightning Fast Shop

Release 0.10.1

We just released LFS 0.10.1. 

Changes

  • Adds Django 1.8 support

Information

You can find more information and help on following locations:

March 11, 2016 07:49 AM


Blended Technologies

New Libraries for UtilityMill

Utility Mill, my web application for building quick Python utilities now has the following libraries available for use in your utilities. Head on over and try it out. Python Excel Tools: xlrd, xlwt, xlutils These tools (Home/Documentation) let you read, write, and manipulate Excel files. They're a ...

March 11, 2016 07:48 AM


Aahz

Python: Call for diversity

The Python community is both incredibly diverse (Python 3.1's release manager was not yet eighteen years old) and incredibly lacking in diversity (none of the regular committers is a woman). I'm now working on creating more diversity in the Python community, and I welcome anyone who wants to help.

March 11, 2016 07:48 AM


Michele Simionato

What's new in plac 0.7

plac is much more than a command-line arguments parser. You can use it to implement interactive interpreters (both on a local machine on a remote server) as well as batch interpreters. It features a doctest-like mode, the ability to launch commands in parallel, and more. And it is easy to use too!

March 11, 2016 07:48 AM


David Goodger

G4G10: Gathering for Gardner 10

An account of my first Gathering for Gardner, a conference for recreational mathematicians, magicians, puzzlers, philosophers, and other curious types.

March 11, 2016 07:48 AM


Katie Cunningham

Announcing: PyCruise 2016!

For years, I’ve been been muttering about doing a Python conference on a cruise ship. As someone who has had to pay out of pocket for many conferences, it made sense to me. Food would be covered, as would lodging. There would be lots of places to hang out, and some interesting ports to see.

This year, I finally pulled the trigger. After some negotiations with Carnival Cruise Lines, I have a contract for PyCruise 2016!

The details

PyCruise will be held from June 20th to June 25th, 2016, and will leave from New York City. We’ll visit Saint John and Halifax, then return to New York City.

Tickets start at $470 per person, and cover lodging, meals, entertainment, and child care.

Rather than having talks, PyCruise will adopt the Open Spaces style of conferences, allowing attendees to focus on collaboration, bonding, and learning new things in a group setting. I’ll open up a CFP for Open Space themes once we have meeting spaces reserved on the ship.

Families are also welcome on PyCruise. In fact, that was one of the reasons why I wanted to have this conference! Non-coding partners will be able to find more than enough to do around the ship, and free daycare gives everyone a break.

Registration

Registration is open! Because we have a set number rooms available, registering early allows you to pick the room that works the best for you (whether that’s the most inexpensive room, or the room with the best location or amenities).

For more details, go to the official PyCruise website, or follow PyCruise on Twitter!

March 11, 2016 07:47 AM


Bojan Mihelac

Django cookie consent application

django-cookie-consent is a reusable application for managing various cookies and visitors consent for their use in Django project.

Features:

Source code and example app are available on GitHub:
https://github.com/bmihelac/django-cookie-consent

Documentation:
https://django-cookie-consent.readthedocs.org/en/latest/

March 11, 2016 07:47 AM


Fredrik Lundh

Webdriver Torso

March 11, 2016 07:44 AM


Hector Garcia

From LIKE to Full-Text Search (part II)

If you missed it, read the first post of this series


What do you do when you need to filter a long list of records for your users?

That was the question we set to answer in a previous post. We saw that, for simple queries, built-in filtering provided by your framework of choice (think Django) is just fine. Most of the time, though, you'll need something more powerful. This is where PostgreSQL's full text search facilities comes in handy.

We also saw that just using to_tsvector and to_tsquery functions goes a long way filtering your records. But what about documents that contain accented characters? What can we do to optimize performance? How do we integrate this with Django?

Hola, Mundo!

We have found that the need to search documents in multiple languages is fairly common. You can query your data using to_tsquery without passing a language configuration name but remember that, under the hood, the text search functions always use one.

The default language is english, but you have to use the right language stemmer according to your document language or you might not get any matches.

If, for example, we search for física in spanish documents that have this word and its variations we would only see exact matches for this query:

=> SELECT text FROM terms WHERE to_tsvector(text) @@ to_tsquery('física');
                                  text                                   
-------------------------------------------------------------------------
 física (aparatos e instrumentos de —)
 física (educación —)
 física (investigación en —)
 rehabilitación física (aparatos de —) para uso médico
 educación física
 conversión de datos y programas informáticos, excepto conversión física
 investigación en física
 terapia física
(8 rows)

And worse, if we search just for fisica, unaccented:

=> SELECT text FROM terms WHERE to_tsvector(text) @@ to_tsquery('fisica');
 text 
------
(0 rows)

To get results that contain a variation of física (físicas, físico, físicamente, etc…) we have to use the right stemmer. Remember that if our stemmer doesn't have a word in its dictionary it won't stem it:

=> SELECT ts_lexize('english_stem', 'programming');
 ts_lexize 
-----------
 {program}
(1 row)

=> SELECT ts_lexize('spanish_stem', 'programming');
   ts_lexize   
---------------
 {programming}
(1 row)

But if we use the right language, it will:

=> SELECT text FROM terms WHERE to_tsvector('spanish', text) @@ to_tsquery('spanish', 'física');
                                      text                                      
--------------------------------------------------------------------------------
 física (aparatos e instrumentos de —)
 ejercicios físicos (aparatos para —)
 entrenamiento físico (aparatos de —)
 físicos (aparatos para ejercicios —)
 física (educación —)
 preparador físico personal [mantenimiento físico] (servicios de —)
 física (investigación en —)
 ejercicio físico (aparatos de —) para uso médico
 rehabilitación física (aparatos de —) para uso médico
 aparatos para ejercicios físicos
 almacenamiento de soportes físicos de datos o documentos electrónicos
 clases de mantenimiento físico
 clubes deportivos [entrenamiento y mantenimiento físico]
 educación física
 conversión de datos o documentos de un soporte físico a un soporte electrónico
 conversión de datos y programas informáticos, excepto conversión física
 investigación en física
 terapia física
(18 rows)

Working with multiple languages

Of course, we don't want to fill in the language manually every time. A straightforward solution would be to store the record's language in its own column. to_tsvector and friends accept a string to set the language but also a column name. So we could write:

=> SELECT text FROM term WHERE to_tsvector(language, text) @@ to_tsquery(language, 'entrena');
                           text                           
----------------------------------------------------------
 entrenamiento físico (aparatos de —)
 clubes deportivos [entrenamiento y mantenimiento físico]
(2 rows)

The only catch here is that the column must be a regconfig. If you are using South or Django migrations to manage your database schema remember to change the type of your language column:

=> ALTER TABLE terms ALTER COLUMN language TYPE regconfig USING language::regconfig;

This way we can query records in their own language.

Even Better: Accented Characters

But what about accented words? If our users don't carefully type them (most users don't) then they won't find anything.

If we search for the word fisica again, even with all the previous setup, nothing shows up:

=> SELECT text FROM terms WHERE to_tsvector(text) @@ to_tsquery('fisica');
 text 
------
(0 rows)

We found this behavior funny (not in a good way). Some words work and some don't. I guess it depends on completeness of the dictionary we choose to use. But we can't depend on the word being on the dictionary. It would be a hit or miss (and, by our experience, more miss) thing.

But don't worry. There is a PostgreSQL extension for that™.

CREATE EXTENSION unaccent;

The unaccent extension can be used to extend the default language configurations to filter out accented characters:

CREATE TEXT SEARCH CONFIGURATION sp ( COPY = spanish );
ALTER TEXT SEARCH CONFIGURATION sp ALTER MAPPING FOR hword, hword_part, word WITH unaccent, spanish_stem;

And, behold! It works great:

=> SELECT text FROM terms WHERE to_tsvector('sp', text) @@ to_tsquery('sp', 'fisica');
                                      text                                      
--------------------------------------------------------------------------------
 física (aparatos e instrumentos de —)
 ejercicios físicos (aparatos para —)
 entrenamiento físico (aparatos de —)
 físicos (aparatos para ejercicios —)
 física (educación —)
 preparador físico personal [mantenimiento físico] (servicios de —)
 física (investigación en —)
 ejercicio físico (aparatos de —) para uso médico
 rehabilitación física (aparatos de —) para uso médico
 aparatos para ejercicios físicos
 almacenamiento de soportes físicos de datos o documentos electrónicos
 clases de mantenimiento físico
 clubes deportivos [entrenamiento y mantenimiento físico]
 educación física
 conversión de datos o documentos de un soporte físico a un soporte electrónico
 conversión de datos y programas informáticos, excepto conversión física
 investigación en física
 terapia física
(18 rows)

Note that you must be superuser to install this extension on a normal installation (not in Heroku Postgres). If PostgreSQL complains about not finding the extension, install the postgresql-contrib package in your distro.

A Note on Performance

To make text search work well in a large dataset, consider creating a GIN index:

CREATE INDEX terms_idx ON terms USING gin(to_tsvector(language, text));

GIN indexes might take longer time to build than GiST indexes, but lookups are a lot faster.

Wrapping Up: Using It in Django

Finally, to make all this easy to use from Python, we wrote a custom queryset and used django-model-utils to embed it on the model's manager:

from django.db import models
from model_utils.managers import PassThroughManager


class TermSearchQuerySet(models.query.QuerySet):
    def search(self, query, raw=False):
        function = "to_tsquery" if raw else "plainto_tsquery"
        search_vector = "language, text"
        ts_query = "%s(language, '%s')" % (function, query)
        where = "to_tsvector(%s) @@ %s" % (search_vector, ts_query)
        return self.extra(where=[where])


class Term(models.Model):
    text = models.CharField(_('text'), max_length=255)
    language = models.CharField(max_length=20, default='sp', editable=False)

    objects = PassThroughManager.for_queryset_class(TermSearchQuerySet)()

Remember our original example on the first part?

>>> Entry.objects.search('man is biting a tail')
[<Entry: Man Bites Dogs Tails>]

March 11, 2016 07:44 AM


Jonathan Street

Lightning talk slides on deep learning with keras

At the DCPython Office Hours event this month I gave a lightning talk on convolutional neural networks implemented with the keras library. The notebook is now up on github.

Deep neural networks are typically too slow to train on CPUs. Instead, GPUs are used. The example in the notebook uses a relatively small network so should be runnable on any hardware.

March 11, 2016 07:42 AM


Ilian Iliev

Django interview questions ...

... and some answers

Well I haven't conducted any interviews recently but this one has been laying in my drafts for a quite while so it is time to take it out of the dust and finish it. As I have said in Python Interview Question and Answers these are basic questions to establish the basic level of the candidates.

  1. Django's request response cycle
    You should be aware of the way Django handles the incoming requests - the execution of the middlewares, the work of the URL dispatcher and what should the views return. It is not necessary to know everything in the tiniest detail but you should be generally aware of the whole picture. For reference you can check "the live of the request" slide from my Introduction to Django presentation.
  2. Middlewares - what they are and how they work
    Middlewares are one of the most important parts in Django. Not only because they are quite powerfull and useful but also because the lack of knowledge about their work can lead to hours of debugging. From my experience the process_request and process_response hooks are the most frequently used and those are the one I always ask for. You should also know their execution order and when their execution starts/stop. For reference check the official docs.
  3. Do you write tests? How?
    Having well tested code benefits you, the project and the rest of the team that is going to maintain/modify your code. Django's built-in test client, Django Webtest, Factory boy, Faker or whatever you use, I would like to hear about it and how you do it. There is not only one way/tool for it so I would like to know what are you using and how. You should show that you understand how important is to test your code and how to do it right.
  4. select_related and prefetch_related
    Django's ORM (as every other) can be both a blessing and a curse. Getting a foreign key property can easily lead you to executing million queries. Using select_related and prefetch_related can come as a saviour so it is important to know how to use them.
  5. Forms validation
    I expect from the candidates that they know how to build forms, to add custom validation for one field and how to validate fields that depend on each other.
  6. Building a REST API with Django
    The most popular choices out there are Django REST Framework and Django Tastypie. Have you used any of them, why an how. What issues have you faced? Here is the place to show that you understand the concept of REST APIs, the correct request/response types and the serialisation of the data.
  7. Templates
    Building website using the built-in template system is often part of your daily tasks. This means that you should understand the template inheritance, how to reuse the block and so on. Having experience with sekizai is a plus.

There are a lot of other things to ask about internationalisation(i18n), localisation(l10n), south/migrations etc. Take a look at the docs and you can see them explained pretty well.

March 11, 2016 07:41 AM